304 épisodes
- Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks!
While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning.
This year we are proud to feature the work of Alex Zhang of MIT.
From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems.
RLMs took over the timeline early this year:
and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra:
and is even today, influencing new research that has more extreme implications than RLMs:
We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie.
We discuss:
* Why AI-generated GPU kernels still leave substantial room for human expertise
* How one expert insight can potentially replace enormous amounts of brute-force token search
* Why PhD students should take research bets that initially look trivial, weird, or pointless
* What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste
* GEV and why a language model does not have to mean an autoregressive text-to-text decoder
* Why Claude Code, Codex, and Pi are structurally more similar than they look
* How harness design can improve compositional generalization across tasks and domains
* RLMs: context offloading, code execution, recursive subagents, and shared memory
* Prime Agent, continual harnesses, and persistent agent-to-agent communication
* Why the model you query in the future may secretly be an entire swarm or scaffold
* OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving
* Why much of an agent swarm may be wasted search — and why convergence is still hard
* Kimi versus OpenAI and different approaches to multi-agent systems
* Open-endedness, Sakana AI, and finding hidden gems in enormous amounts of generated work
* Why current frontier models may already have a large capability overhang
* Speculative programmatic tool calling and overlapping tool execution with generation
* Whether English, code, or an entirely new “Neuralese” constrains how models reason
* AI for science, fast-moving benchmarks, and how Alex chooses what research problems to bet on
Alex Zhang
* Website: alexzhang13.github.io
* X: @a1zhang
Timestamps
00:00:00 Introduction
00:00:49 GPU Mode, KernelBench, and AI-Written Kernels
00:07:38 Human Expertise vs. Brute-Force AI Search
00:13:20 Research Taste and Taking Big Bets
00:19:28 GEV and Rethinking the Language Model
00:29:03 Video Game Agents and the Harness Problem
00:31:01 Why Claude Code, Codex, and Pi Are So Similar
00:36:42 Harnesses as Compositional Generalizers
00:44:24 RLMs Explained
00:52:01 Prime Agent and Persistent Subagents
00:57:41 RLMs in the Wild
01:00:30 OpenAI Swarms and the Future of Language Models
01:07:26 Open-Endedness and Sakana AI
01:15:52 Kimi vs. OpenAI Agent Swarms
01:20:06 Capability Overhang and Speculative Tool Calling
01:28:19 Neuralese, Future Research, and AI for Science
Transcript
Introduction: Alex Zhang, RLMs, and GPU Mode
Swyx [00:00:00]: All right, we’re here in the studio with Alex Zhang, I guess most famously of, RLMs, but you have a few other affiliations. Welcome to the show.
Alex Zhang [00:00:12]: Yeah. Thank you for having me.
Swyx [00:00:13]: Yeah. I guess GPU Mode as well?
Alex Zhang [00:00:15]: Yes, GPU Mode as well.
Swyx [00:00:16]: You were shepherded in by Mark Saroufim. Not everyone gets that kind of welcome.
Alex Zhang [00:00:19]: Yep. Yeah. Yeah. I’m very close to all the people in GPU Mode, so yeah
Swyx [00:00:23]: Yeah
Alex Zhang [00:00:23]: We often end up working together in various capacities, like even beyond just GPU Mode itself, so.
Swyx [00:00:29]: Yeah. Can we explain, so people who are not that close Don’t know about this. It’s just a-- it’s a Discord. It used to be focused on, I guess, CUDA Mode, and then generalized a little bit. it was started by Mark.
Alex Zhang [00:00:41]: Yep.
Swyx [00:00:42]: It was basically like. To me, it’s like the hiring pipeline of the PyTorch team.
Swyx [00:00:45]: And then you left PyTorch.
Alex Zhang [00:00:47]: Yep.
Alex Zhang [00:00:49]: Yeah. So it used to be, I think, it started actually around when I was in college, in like 2023. I think it was started by Mark, Andreas, and Jeremy Howard. The original premise was just like, it was a GPU-- or it was a Discord dedicated to learning how to write GPU kernels, and they had, like, lectures. That was basically the extent of it. and I got interested in it because I was writing GPU kernels. It was actually out of, like, pure chance. I was interning at Snapchat at the time, and
From CUDA Mode to GPU Mode
Swyx [00:01:22]: Rexis.
Alex Zhang [00:01:23]: Yeah, I was very bored with Rexis. So, they had a project where, like, they were interested in writing. It was this paper called Infinite Attention. It was like a Google paper.
Swyx [00:01:35]: Yes, we’ve covered it on Paper Club.
Alex Zhang [00:01:36]: Yes, yeah. So I was interested in whether or not you could write specialized kernels for it at Snapchat. it didn’t. Nothing really came of it, but I joined GPU Mode. At the time, it was called CUDA Mode, I think for, like, legal reasons or something, they changed the name. But I met Mark, I met Matei, I met a bunch of other people that were very involved in the community. And then Mark had pitched this idea called Popcorn, which was now what you see as the leaderboard today. But the general idea was like. I think all of us had this, like, intuition that GPU programming is, like, very similar to if you guys have done, like, competitive programming. It’s a, it’s a not. I don’t mean to say, like, they’re transferable skills.
Popcorn, KernelBench, and Automating GPU Kernels
Swyx [00:02:17]: You have constraints. You code golf a little bit.
Alex Zhang [00:02:19]: Yep.
Swyx [00:02:19]: Yeah.
Alex Zhang [00:02:19]: Yeah. And there’s like. There’s actually a surprisingly small space of optimizations that people do. and there’s actually not that many kernels per se that people are interested in optimizing. And so we kind of had this thought that, like, if you had enough data, like, in the same way that Codeforces, there’s millions of problems. If you could do this with GPU code, like, you could scale and automate kind of GPU kernel development, which for researchers is a huge deal. ‘Cause I think one of the bigger bottlenecks. Like, if you look at, like, Mamba, for example, like, they release the paper with kernels because otherwise, like, you can’t really use it in any meaningful way? And not everyone has, like, a Tri Dao on their team. So we’re very interested in this. KernelBench kind of spawned from that too, of like, can we get LLMs to automate, GPU kernel code? And I think that was like a. It was a very fun time. It was like between college and my PhD, and yeah, I had a really pleasant time doing stuff with GPU Mode. Now I kind of just help with the lectures sometimes. I’m not as involved, and I think in general, like, we don’t have as many competitions as we used to. But, yeah, I still keep in touch a lot with everyone there.
Swyx [00:03:27]: Is there a friendly rivalry? Because I think the previous community that used to do this was like MLSys, MLPerf
Alex Zhang [00:03:32]: Yeah
Swyx [00:03:32]: Kind of thing. Is there a friendly rivalry? Is this like just new generation MLPerf, or what’s going on?
Alex Zhang [00:03:38]: The nice thing about GPU Mode is that it is also a community in the sense that, like, a lot of the lectures are very easy enough for a beginner to follow and ask questions and things like that. And, like, the competitions are, like, somewhat, not secondary, but, like, you can participate in them to learn. I think with a lot of, like, MLSys, MLPerf kind of benchmarks, like, for the most part, like, only serious labs and companies participate. Like, seriously in them, at least. That was my understanding of it. I could be wrong. But I think also beyond GPU Mode now, one thing that has been really exciting is there’s a lot more websites and, like, people that work on hosting competitions. Like, I think there’s this.
Alex Zhang [00:04:21]: I think there’s this website called, like, LeetGPU or something, and it’s like leet code for GPU problems.
Swyx [00:04:26]: Wow.
Alex Zhang [00:04:26]: There’s, like, other ones that I’ve. Like, we’ve, we’ve seen. Like, there’s many that have kind of spawned and, like, talked on GPU Mode, and like, it’s very exciting in the sense that I think GPU programming used to be super niche, like when I was interested in it. And the only reason I got interested in it was Tri Dao gave a talk at Princeton because he was, applying for faculty. he is faculty there now, but I listened to his talk on FlashAttention in like 2023, and I was like, “Wow, this is like the coolest thing ever.”
Alex Zhang [00:04:56]: And I was like, “This is like. This is what everyone should be working on.” I guess, like, vLLM and stuff had come out too, and it was like, “Oh, we should be writing kernels.” But now it’s like, everyone writes kernels. Like, everyone. It’s, it’s. I think it’s actually almost saturated in some sense, as a field.
AI-Written Kernels and the Verification Gap
Vibhu [00:05:11]: Any interesting takes for people that wanna get into it? So I think one of the biggest news is GPT-5.6 Wrote more efficient kernels, so Terra and, Luna could be 80% cheaper.
Alex Zhang [00:05:24]: Yeah.
Vibhu [00:05:25]: And then we’ve seen other competitions where people are, like, setting records, and they’re like, “We’re doing some auto research loop,” and these are people that don’t have a background
Alex Zhang [00:05:34]: Yep
Vibhu [00:05:34]: In any kernel writing, right?
Alex Zhang [00:05:36]: Yeah. So even on the GPU Mode leaderboard, if you look at, like, a lot of the recent problems, almost all the solutions are AI generated. However, you’ll notice on the leaderboard. So there’s this guy named Gauners who’s, like, a very, like, regular member of GPU Mode. We’ve always known for a long time that he’s, like, a super cracked, like, GPU kernel writer. One thing we discovered on this leaderboard is, like, almost ev-- Like, he also used AI to help him with these solutions, but for the most part, like, he helped prompt and move it in certain directions. we found that, like, his kernel was, like, basically the only one in, like, the top 10 that was actually stable in, like, actual, like. end-to-end systems. And it does bring into question, like, it’s not. Like, GPU kernels have a verification problem. Like, we’ve kind of known this. It’s been a problem since KernelBench was released. Like, there’s a lot of reward hacking that goes on. But you al-- Yeah, you also notice, like, the lines of code is a lot smaller, but
Vibhu [00:06:33]: Yeah, I was gonna ask, is that noticeable, or is it just
Alex Zhang [00:06:36]: Yeah, no, it’s, it’s
Vibhu [00:06:36]: Okay
Alex Zhang [00:06:36]: It’s definitely, like, very important, and I think, like, it’s, it’s really interesting that still there’s a lot of alpha in being good at writing GPU kernels.
Vibhu [00:06:44]: Okay, so there is a gap from verifying
Alex Zhang [00:06:45]: There definitely is, yeah. I think, like. And this applies to a lot of AI systems as well. Like, I think, even with the most recent, like, math proofs and stuff, like, it doesn’t necessarily mean mathematicians are obsolete. these companies still hire mathematicians, like, to do, whether it be, like, data labeling work or even just, like, steering the models to solve problems. Like, there is still a lot of alpha in being, like, knowledgeable in these things, so.
Swyx [00:07:12]: Is it just knowledge, or is it also there is just more planning, and is there, an emergent style of planning that works better?
Alex Zhang [00:07:22]: I think it’s, it’s a mix of. Maybe this is what you mean, like intuition for
Swyx [00:07:27]: Something like that
Alex Zhang [00:07:28]: How to solve the problems.
Swyx [00:07:29]: Like, for example, I always diagram my code.
Alex Zhang [00:07:31]: Yeah.
Swyx [00:07:31]: Right?
Alex Zhang [00:07:31]: Yeah.
Swyx [00:07:31]: And then, like, if there’s a part of the diagram I don’t understand, I work until I understand it. Otherwise, I, it’s not allowed.
Alex Zhang [00:07:37]: Yeah.
Swyx [00:07:38]: Yeah.
Alex Zhang [00:07:38]: So I think it’s, like, it’s a mix of those things of, like, the people who work. Like, the people who know how to look at these problems and how to solve them, like, also know how to use AI to do them. Because, like, you’re acting as a very strong verifier. Like, if you are knowled- or if what to do and you are also. Like, I think the thing that we’ve kind of discovered with all these agent swarms and things like this is, like, when you throw enough compute at a problem, you, like, can sufficiently explore solutions to that problem. But oftentimes, like, maybe you can burn, like, 100 billion or a trillion tokens on something, but if you bring in someone who knows something about the problem, they can uncover something for the model that would, like, erase that one trillion token spend. it’s not, it’s not super clear, like, what exactly the trends are here. But I think, like, there are so many problems in the wild still right now that we want to solve, and, like, we can’t afford to just always, throw as much compute as possible at it. Like, there is still an efficiency aspect of all of these things that is super important.
Speed-of-Light Limits, Memory, and Megakernels
Swyx [00:08:41]: Is there, like, a theoretical right answer that you just calculate based on physics, and then you just get close to the physics limit?
Alex Zhang [00:08:49]: Yes. So for GPU kernels, you can compute. It’s actually not that easy to compute sometimes, like, depending on how complex the problem is. Like, for matrix multiplication, it’s very easy to compute, this, like, speed-of-light kind of, estimate of what the fastest kernel can be. And, like, this is also assuming, like, maybe all of your, all your data starts on the CPU, or maybe it starts in DRAM, on the GPU, et cetera. Like, this changes these numbers slightly, but
Swyx [00:09:18]: The transfers and all these things, yeah.
Alex Zhang [00:09:20]: Yeah. But I will say, like, it’s not clear, though, like, in a lot of cases if it’s even possible to hit this theoretical number, if that makes sense. Like, this is assuming, like, perfect overlapping and transfer of data, and, like, there’s maybe some bottleneck that you can’t get around. But often, the kernels are not even close. Like, that we write are not nearly close enough to this number to be, like, meaningful at all.
Swyx [00:09:43]: Yeah. And is it speed that matters? Do you also care about, obviously memory, which
Alex Zhang [00:09:48]: Mm
Swyx [00:09:48]: Feeds into speed? Do you care about power consumption? So one of my, one of our top pods of the year was Geoff Dean, who was like, “Actually, I just tracked the microjoules or, like, the nanojoules, picojoules.”
Alex Zhang [00:09:59]: Yeah, it’s often picojoules today.
Swyx [00:10:01]: Picojoules.
Alex Zhang [00:10:01]: Yeah.
Swyx [00:10:01]: Do you care about that?
Alex Zhang [00:10:03]: So I don’t.
Alex Zhang [00:10:04]: Yeah. I guess maybe I’m not, I’m not as
Swyx [00:10:06]: But everything here is speed, right? Like
Alex Zhang [00:10:07]: Everything here is speed
Swyx [00:10:08]: Nobody’s counting picojoules.
Alex Zhang [00:10:09]: But I-- There’s a caveat here, which is, I think, like, there is speed in the context of a single kernel, and there is speed in the context of a larger problem, like maybe the N10 model. Because, like, one thing to consider, and this is why it’s important to talk about what speed-of-light is referring to, because in these cases for the kernels, like, we always start with everything in, like HBM, for example, right? But you can imagine that, like, an end-to-end like an end-to-end model, what you might wanna do between two layers is, like, you might sacrifice the speed of the first operation to keep things in the cache for the second operation. And, like, these are things that, like, you can’t really get out of, in isolation with, like, these kinds of kernels. And, people call this, like, the fusion or, like, the
Swyx [00:11:01]: Megakernel
Alex Zhang [00:11:02]: Kernel fusion problem. Yeah, or, like, megakernel stuff. And it generally only applies, like, when you are, like, memory-bound in most cases. But this is something that, like, also there is this question of, like, as these models get better, like, should we just be generating like, megakernels? Is that, like, what we want?
Vibhu [00:11:20]: What’s, what’s your take?
Alex Zhang [00:11:21]: I think that this is really difficult because you need the data to do this. And I think, like, I have yet to see an example in the wild of, like, we bootstrap the ability to solve a very difficult class of problems without any examples. and I think, like, the other reason why I think maybe this isn’t that interesting is that at the level of an individual kernel, A, like, they’re not that, they’re not as complex, but B, you’re somewhat confident that there’s not as much structure in a single kernel. But, like, in a megakernel, like, I would be more inclined to believe that, like, a compiler would be better here. Like, some compiler over, like, higher level- Ops makes sense, because in general, like actually, I think mega kernels are very, like the pieces are very composable of like the individual kernels. There’s some areas where you might wanna do like weird fusions and everything, but in general, I think these are cases that like a compiler can probably handle. And there is a company that’s working on this from what I understand that has given some talks on GPU mode as well.
Swyx [00:12:28]: Yeah, I wanna basically cluster all the GPU mode discussions here because obviously there’s other parts
Alex Zhang [00:12:32]: Right.
Swyx [00:12:32]: That we need to move on to.
Vibhu [00:12:33]: I think there is something to plug. You guys do host a lot of really good lectures. They’re all on YouTube. People can follow. And you
Alex Zhang [00:12:39]: Yes
Vibhu [00:12:39]: Lead quite a bit of it. You’re still quite involved.
Alex Zhang [00:12:41]: I used to. sometimes I still do. I think they’re mostly Mark. Mark is the one who usually does them. Matei does sometimes as well, but, yeah, I highly recommend them. They are extremely good resources. Like, I think it’s kind of crazy how much people share on there, so yeah.
Swyx [00:12:59]: ‘Cause like if you’re there, like you’re very, like you’re exactly the right audience?
Alex Zhang [00:13:03]: Yes, exactly.
Swyx [00:13:03]: Like this isn’t gonna reach the mainstream.
Alex Zhang [00:13:04]: And there’s a lot of like introductory material as well, that we’ve put on, that I think is useful for people.
Benchmarks, Princeton, and Research Taste
Swyx [00:13:10]: KernelBench was kind of influential. I just wanna see like, that was last year.
Alex Zhang [00:13:13]: Yep.
Swyx [00:13:14]: What other ongoing work do you wanna shout out that people should pay attention to? ‘Cause obviously you’re involved in this
Alex Zhang [00:13:20]: Yep
Swyx [00:13:20]: Field.
Alex Zhang [00:13:20]: Yeah. I will give maybe the background story of like I am actually involved in a lot of benchmarks, or I used to be, maybe prior to my PhD. It started because I was at Princeton. I worked with the SWE-bench team there.
Swyx [00:13:33]: John, Carlos
Alex Zhang [00:13:33]: John, Carlos
Swyx [00:13:34]: Ofir
Alex Zhang [00:13:34]: And Ofir. They’re all great. Like
Swyx [00:13:36]: Karthik
Alex Zhang [00:13:36]: I love them. Yeah.
Swyx [00:13:37]: There’s basically this, like I think people don’t understand how much benchmarks come from the same group
Alex Zhang [00:13:42]: Yeah
Swyx [00:13:42]: At Princeton.
Alex Zhang [00:13:44]: It is crazy.
Swyx [00:13:45]: Do Xun Yu?
Alex Zhang [00:13:45]: Yes. Yeah.
Swyx [00:13:46]: We had him on a pod before. Now he’s like running Tencent.
Alex Zhang [00:13:48]: Yeah, now he’s like, he’s like a superstar. when I met him, so he was advising my friend Michael Tang, who is now at Anthropic, but they worked together a lot. We were like the two undergrads in Karthik’s lab. I-- And then some others joined later as well. But yeah, Xun Yu is great. I did not know he was like such a superstar until like later on, like after I left, but
Swyx [00:14:12]: Yeah. like, okay, so there are very few PhD students. Like yours is like the next one. Like once a year, we feature someone like who is like basically entire PhD, has been like on target.
Swyx [00:14:25]: There’s not that many of them. Xun Yu was like clearly one of them. And, Jack Morris is another one. And like, we talked, before the show, we talked about research taste.
Swyx [00:14:33]: Right? Like somehow some grad students just have a very blessed career where like, yeah, mostly like, yep, this is like going to stick around, relevant, everyone should know this.
Alex Zhang [00:14:42]: Yeah.
Swyx [00:14:42]: And then others, just nothing.
Alex Zhang [00:14:44]: I think this is also true of like even people within like industry labs as well. I think it’s just like grad students are a lot more visible. So you just see, like you see, like, there are some people who
Swyx [00:14:56]: Yeah, you can publish
Alex Zhang [00:14:57]: Who really like get lucky and like, or it’s, it’s a mix of being lucky and also being very smart and things like that. I think like with research taste as well, like I think it gets developed through opportunities, at least in my case. Like I got-- I was very fortunate to have like taken the path that I took, like working at Princeton and then like later, like finding my like Omar at MIT. Like he’s a fantastic advisor. I will say, though, I find that the most successful research from grad students or like in academia comes when people care about problems that maybe like most people in industry are not looking at. I think this is the issue that like a lot of grad students work on things that benefit, like that look good to an industry lab. Like for example, they’ll work on some, like some benchmark that’s really popular now. I think benchmarks in 2023 were a very different story than benchmarks now. There are a lot of people that work on like harnesses and like meta-harnesses and like specific harnesses for XYZ task. And when you really think about it, the reason someone would work on this is like maybe there’s like a clear goal shaped around the models that we have today of like, this is what I want to see. But like, I’ll give, I’ll give the like RLM, like the recursive language model paper as an example, because I think like it’s a super simple idea. I think when it came out as well, like there were a lot of people that were like, when they see something like that, they’re like, “What is even the purpose of this?”
Swyx [00:16:27]: Or like too cool.
Alex Zhang [00:16:28]: Yeah. Like why, like what? This is just subagents or something, right?
Alex Zhang [00:16:32]: And I think it’s like when you get a reaction like that, it’s almost like a good sign in the sense that like it’s clear that people aren’t thinking about what the purpose of this is. And I will give another example of like SWE-bench. When SWE-bench came out, Ofir loves to tell this story. When it came out, like nobody cared. Like everybody was like, “This is an impossible task. Like why would we ever even consider this as a benchmark?” And it wasn’t until Devin came out that everyone was like, “Whoa, like this is something we wanna hill climb.” And I think this is, this rings true for. You tend to see that a lot of ideas. I think like the. My favorite, I guess, example of this is Eric Seligman’s work, with like STaR and like Quiet-STaR. Like I think when you read the paper, at least when I first read the paper, I was like, “Is this not like an obvious idea?” Or maybe not. I don’t know. I was like, “Oh, this seems really simple.” Or like chain of thought, and the same thing. Or like Xun Yu’s react. It’s like, okay, like, yeah, sure. But then like when you really think about it’s like why. What is the value of the paper? And I think it comes from like, it tells a bit of a story as to like what you want the field to look like. And that is something that it’s very hard to do this in academia because if you look at all these papers, Quiet-STaR, ReAct, RLMs, SWE-bench, none of these papers are. It’s not like a GPT-6 Astro release? It’s not like everyone’s like, “Oh my gosh, like I’m gonna use this now and this is the best thing in the world.” Like academia just can’t afford to do this, at least right now. I. There’s a whole slew of reasons why I think that should change, but I think it’s like. If you don’t have. As a PhD student, I think you’re in such a unique position where you can work on literally whatever you want for the most part. If you’re not taking advantage of that and working on things that, like, nobody cares about or, like, people see as some trivial thing, like, “Oh, I thought about this, but, like, I don’t use it,” I just think, like, in the end, the research is just never gonna be that interesting because you kind of need to take big bets if you’re gonna be in academia. Because otherwise, I think, like, just go to an industry lab. Like, they have tons of resources, tons of talent. Why constrain yourself in an area where you don’t have a lot of resources and, like, there’s not even that many people around? And I think it’s just. it literally just comes down to, like, big bets. like, you just have to take big bets, and, like, a lot of them will fail? Like, that’s just. it’s, it’s natural. But I think
Alex Zhang [00:18:55]: That is, as a PhD student, like, that’s the biggest advantage you have over any single person at another lab because you don’t have to deal with bureaucracy and all these other things.
Swyx [00:19:07]: Fair enough.
Alex Zhang [00:19:07]: Yeah.
Swyx [00:19:08]: I ask a lot of people this question, and usually they hand-wave away. So I think. I appreciate that you’re actually giving a thoughtful response on Like, no, like, this is your unfair advantage because everything else is biased against you, basically.
Jev and Breaking the Autoregressive Decoder Paradigm
Alex Zhang [00:19:20]: Yeah, exactly. And so, like, it’s honestly. I will bring up Jev as an example because it’s, it’s not an academic
Swyx [00:19:26]: Wow, okay.
Alex Zhang [00:19:27]: It’s not an academic project.
Swyx [00:19:28]: Yes.
Alex Zhang [00:19:28]: I want to bring this up because this also happened with RLMs and, it happens with many other works. Like, things get overhyped, right? To an extent, like, something gets overhyped and then people are like, “Why is this overhyped?” Like, “This is trivial. This is stupid.” And I saw the same thing with Jev because I think the release was like. there is this whole thing about, like, academics, or they’re not an academic group, but, like, people have to do branding and they have to, like, kind of market their research. And so, like, I understand, but I think there was a lot of discourse about Jev just being, like, something we’ve known for years. And I think it’s kind of missing the point of, like, why is such a system so interesting? It’s why is it not just some stupid NLP classifier that, like, we’ve, we’ve been doing, back in our intro ML classes or something? I think what’s really interesting about Jev is that it kind of opens up this question of, are language models correct? Like, in the form that they’re in, can we consider a different design space other than text-to-text? Because what they’re doing is they’re basically saying like, “I will take advantage of this language model backbone. Like, I know it captures a lot of information about language, but I’m going to change the output space of the model to give you a trade-off, which is I will do very fast inference over.” Like, if you have some prior about this problem, like, let’s say I only need to make a binary classification. Am I gonna ask my language model to do this and pay, like, a 400X cost? Like, no. That’s-- it’s, like, silly, right? and I think for the longest time, because the labs are the only places that control, you’re never gonna use something other than, like, GPT-4 or GPT-6 or Fable because they’re the best models. But because of that, like, people have gotten kind of accustomed to this idea that a language model is just a autoregressive decoder. Like, we have accepted this. And I think when RLMs came out, it was the same thing. Like, one of the comments, like a very frequent criticism I got was like, “This is not a language model.” Or like, “When I look at this, like, I thought it was a new architecture, but it’s actually not.” And my response to that is like, “Well, a language model is just modeling language. It doesn’t have to be this transformer decoder,”? and Jev is really interesting in that, like, we now have a new
Alex Zhang [00:21:49]: Thing to tune, which is like, what is the output space and how does this affect inference latency? and I think we can actually start asking this about various parts of the language model itself. we are seeing this too with, like, loop transformers. It’s a similar idea of a lot of the attention around it was like, “This is a silly idea.” Like, “Why? Who cares about this?” But it’s like, it is a simple idea, but it’s actually. it opens up a whole new set of questions that I think, like, especially if you’re a PhD student, these are the things that you wanna answer. Because I think it’s like we don’t know. For Jev, for example, we don’t know how far we can take this. for loop transformers, we also don’t know how far we can take this. What if you loop only a subset of the model? what if you route to, like, only. like you have some router to different parts of the model? Like, can you mimic what you do in a harness inside of the model? What can you bridge between a harness choice and a model architecture choice? These are all questions I think that open up with works like this, and that’s, like, really exciting. Jev in particular, when I saw it, I was like, “This is actually really useful for RLMs.” Like, I think it’s, it’s. it makes sense ‘cause the biggest bottleneck in RLMs or swarms or systems like these is they’re slow. When you do multiple language model calls all the time, you’re not distributing your compute correctly because, like, maybe there’s something trivial that you just want a simple model to do, but you can’t do it because your language model is just this bulky thing? So I’m very excited. I think we will start to see new types of models emerge beyond just the bog-standard frontier model, and that is like. there’s so many things that you can do with these, like, new trade-offs.
Swyx [00:23:34]: I’ll also shout out Thinky with their interaction models.
Alex Zhang [00:23:36]: Yes. Yeah. Another great example.
Swyx [00:23:38]: Yeah. So, like, basically try to break the paradigm from sequence to sequence, decoder only, and just literally do anything else.
Alex Zhang [00:23:46]: Yeah.
Vibhu [00:23:46]: I think there’s, there’s a level of if you’re trying to compete, you’re not gonna compete with a Frontier lab doing an autoregressive
Alex Zhang [00:23:54]: No.
Vibhu [00:23:54]: Decoder on. Like, the amount of compute scaling resources they Even Thinky will not. okay, there may be one of the handful that can, but you’re not really gonna do much in that at least.
Alex Zhang [00:24:06]: I don’t know too much about Thinky, or I don’t wanna say anything either, but it’s like if their strategy is just to replicate OpenAI or Anthropic, like that’s a horrible strategy.
Alex Zhang [00:24:15]: Because, well, because, like, they just don’t have. Like, you kinda just have to think of it in terms of, like, what advantage do you have? And if you’re going to use the same setup. I’m sure they’re not, but it’s like if you’re going to do the same setup, like you’re basically competing on the things you can control, which is data and compute, and obviously they cannot compete with the Frontier Labs on that. So yeah, it makes sense that, like, if you are a neo lab, like. Actually, I don’t know if you would consider them to be a neo lab, but I guess, like
Vibhu [00:24:40]: Yeah. That’s why they’re, they’re in there.
Alex Zhang [00:24:42]: I guess they’re kind of a weird one, yeah.
Vibhu [00:24:43]: They’re in their list. They shipped Inkling. Like, they can’t.
Alex Zhang [00:24:46]: Anything other than OpenAI or Anthropic, maybe like Meta and GDM, like you just, you gotta do something else? Like, it just. It’s the sad reality, but I think. I actually think it’s a good thing. I’m very happy that, like, scale works and these companies will just keep doing it because it opens up, like, potentially new players, like if they uncover something really interesting. Because I sort of have my doubts that this is, like, seriously going on at Frontier Labs, ‘cause it’s like why would you do that? like, why would you take the risk of allocating a large amount of compute to new bets when the old bet already works? So
Vibhu [00:25:24]: And I think that’s what spins off a lot of neo labs, right?
Alex Zhang [00:25:27]: Yeah.
Vibhu [00:25:27]: You have a side bet and you don’t get compute, and you’re like
Alex Zhang [00:25:30]: Yeah
Vibhu [00:25:30]: “Okay, I’ll go, I’ll go do that.”
Alex Zhang [00:25:31]: Exactly.
Vibhu [00:25:31]: And, your example of the potential upside is something like Jev, which is X hundred times cheaper, comes out, and maybe it is language model.
Calibration, Fast Classification, and New Model Trade-offs
Alex Zhang [00:25:41]: Yeah.
Vibhu [00:25:41]: In this case, it’s just different.
Alex Zhang [00:25:42]: Yeah.
Swyx [00:25:43]: Yeah.
Swyx [00:25:44]: So no speculation on what Jev actually is?
Alex Zhang [00:25:46]: I guess I have some guesses for what it might be. I have seen some people say like, “Oh, it’s like a diffusion thing.” I guess that
Swyx [00:25:56]: Which is the parallel decode, right?
Alex Zhang [00:25:58]: Yeah, parallel decode. Honestly, I think regardless of what it actually is, ‘cause I think you can. I’ve seen some, like, open-source replications of it. What is really exciting to me about what they did is I’m not entirely sure what their optimization objective was and how they trained it. And I think, like, this is a thing for RLMs that we’ve also been thinking about, which is like, okay, like RLMs are a very simple idea. If I come out with this paper, like anyone can use it now. But what distinguishes The actual value of an RLM is whether or not you can train it properly, and whether or not maybe you can mold some architecture around the system to make it really good. And that’s something that, like, I’m actively working on, I guess. But I think for them, like, they figured out a way to train the system, which is completely non-trivial. Like, I actually don’t really know how they did it. And I’ve seen some comparisons online of, some people are claiming they used Qwen, or they post-trained on top of Qwen, but every open-source Qwen that you use is gonna be worse, ‘cause whatever they did to train it clearly works very well. And so that’s, that’s very exciting.
Swyx [00:27:04]: There’s one element of calibration Which, is a rare topic that I don’t think people even knew about or understood. We covered it with, our conversation with Clementine Foley of Hugging Face, and she used to run the evals, at Hugging Face, which is basically the idea that, models are attuned to give you the most likely next token. But, they’re gonna lie to you when you ask them, “How confident are you?” Because they’re just gonna give you the most likely next answer instead of, like, actually, like, no, let’s calibrate. Like, I am actually fifty percent sure, or I am twenty percent sure, and, like, let’s try to calibrate that. I would say, like, if anything, I think that actually that’s pretty easy to generate synthetic data around Because you can sort of see the truth and then synthetically generate a bunch of answers, have Jev classify it, and then compare with ground truth.
Alex Zhang [00:27:52]: Oh, I see.
Swyx [00:27:52]: That would be my reverse engineering of this.
Alex Zhang [00:27:54]: Yeah.
Swyx [00:27:54]: I’ve actually. I think calibration is probably the under. Like, people are just using it as a very fast classifier But they’re actually not even using the probability or calibration estimates.
Alex Zhang [00:28:04]: Yeah.
Vibhu [00:28:04]: I think it’s also still just misunderstood to reiterate. When you ask a model, “How confident are you?” it will spew out what, forty-three percent. the big delta is this is a grounded classification, right?
Alex Zhang [00:28:16]: Yeah. Yeah, I’m, I’m very excited to see what people do with this model. is it gonna solve everything? Like, no, of course not. But I think it solves a class of problems that we traditionally struggled with, which is low-latency things. So I love the examples with games. That’s actually like. I
Swyx [00:28:34]: Yeah, the Doom example
Alex Zhang [00:28:35]: Yeah
Swyx [00:28:35]: Was very good.
Alex Zhang [00:28:35]: I have a benchmark on language models playing video games. I’ve always been fascinated by whether you can have an intelligent system play new games and things of this nature. And so I think it’s really cool that they have sort of a unique way to do this, to capture language and understanding in, like, a fast, a very fast model.
Language Models Playing Video Games
Vibhu [00:28:59]: Oh, while you’re on the topic, anything you wanna point out for video games?
Alex Zhang [00:29:02]: Oh, yeah.
Vibhu [00:29:03]: This is. you did do a benchmark on any project, right?
Alex Zhang [00:29:05]: So, yeah. I guess these numbers are very outdated
Vibhu [00:29:08]: Yeah.
Alex Zhang [00:29:08]: Because a lot of the models are very different. And I’ve seen actually people run. There are some folks out there that are running newer models on these games, which is really cool. I guess the general premise of this benchmark was we just want to see if vision language models are good enough at just, like, plugging into games with the latency constraint included. ‘Cause this actually. I came out with this right after Claude Plays Pokémon came out.
Alex Zhang [00:29:33]: So this was, like, two years ago, which I guess is, like, ancient now. But I think what’s really cool about this suite of tasks, it’s very diverse in terms of what games they are. And also, I think most of the games are games that people know or, like, have seen before. I saw, yeah, Jeff playing Doom. I will say I don’t think. I think they were just playing, like, really simple levels and stuff. But honestly, like, most models still can’t really do. Or I don’t actually think any models can solve these games very meaningfully. Like, there are some games that they can. I think I’ve seen Astra be able to solve the Kirby game. And we also, for this benchmark, we intentionally designed a really minimal harness. And I will get to this point about harnesses because I think there’s a whole conversation to be had about, like, what is the value of a harness? Like, what is even the purpose of a harness for a model? But in general, like, I think it’s. yeah, I hope to see very quickly or very soon, like, all of these games beaten by newer models.
Vibhu [00:30:35]: Yeah, it’s interesting. Like, the old Cloud Place Pokemon, they, like, read state from RAM and saw what tiles are walkable and whatnot. We did a podcast with them A long time ago.
Alex Zhang [00:30:45]: Gotcha.
Vibhu [00:30:46]: Yeah. Just fun.
Swyx [00:30:46]: Yeah, and it’s similar. Like, Jeff doesn’t have vision
Alex Zhang [00:30:48]: Yes
Swyx [00:30:48]: So you have to kind of feed in,
Harnesses as Compositional Generalizers
Alex Zhang [00:30:50]: Yeah
Swyx [00:30:50]: These, like, game state and all these things. let’s go right into the harness stuff
Alex Zhang [00:30:54]: Awesome
Swyx [00:30:54]: Because you brought it up. Language model harnesses are compositional generalizers.
Alex Zhang [00:30:59]: Yes.
Vibhu [00:31:00]: You struggled to read that one.
Alex Zhang [00:31:01]: Explain. Yes. Okay. So I have been a little unsatisfied maybe with how people think about harnesses, because people compare like, “Oh, like, I love Claude Code, I love Codex, I love Pi.” Like, “No, I love Oh My Pi, I love Prime Agent.” To be honest, I think all of them are the same. Most of the design decisions or, like, the design choices around these harnesses are the same. Maybe Prime Agent is a little bit different because it’s, like, inherently an RLM. But in general, like, I think we can be a lot more creative with harnesses. And what by that is if we think about this from the perspective of what exactly is the harness doing for the model? Well, basically, when you’re trying to solve a problem and you want to use a language model to solve it, like a very difficult task, one thing that we have discovered is that next token prediction is a really awkward form to do a lot of these tasks. So for example, take SWE-bench. When you’re navigating a code base, like, are you going to be able to figure out how to do all of this with a single language model call? Like, you just say, “Solve code,” or like, “Solve my query over this code base.” No. So we rely on a harness to help you do these kinds of things. And I think what is interesting is, like, a harness is a very opinionated program over how you want a language model to be form fit over a problem. And I bring this up because when we think about the, what a harness is doing, we should really think about, like, what choices in the harness let the language model solve this task? And can I actually just have a language model that just does this? Because a harness, if you think about it, now that loop transformers are a thing, I think what’s really interesting about it is you can model a looped transformer in some ways, like, with a harness as well, right? You’re just looping over the model. Now you can say like, “Oh, I’m not decoding,” so it’s, like, a little bit different. But in general, we, for whatever reason, have stuck with the same model architecture choice forever. And I. And there’s many arguments for why, but clearly, like, we are now training models around harness rollouts. And so there is this very awkward way of doing training over harnesses, which is that we train a language model to act within a harness, but, like, now it’s like a really long, maybe, like, multiple agent rollout that we’re doing. And there’s, like, really awkward, hacky ways of doing this. So what this blog talks about is like, well, one way you can think about what is going on here is if the harness is basically helping the model solve a particular task, can different harness design choices actually do something a little bit more meaningful beyond just, “Here are some tool calls that will help you. Here is a way to grep through your code base.” And so this actually. The idea for this blog came with the RLM idea as well. We just didn’t package it that way. And I think this is actually true. There are many other ideas around RLMs that, like, we will be coming out with, but were all there from the beginning. these are all design decisions around. I think with what is. What I like about the RLM is that there were many iterations and versions of different abstractions that I was interested in doing, and ultimately the RLM made the most sense. But there’s a lot of reasons that aren’t public as to why that’s the case. you will see. So in this example, one of the things that we see with an RLM is that if you sufficiently offload context and write. ask the model to write code over that context, you get this really weird but useful property, which is that when the model recognizes during training how to solve a task, it turns out that the solution to many tasks is very similar across, like, tasks where you don’t even. Like, it’s not even that clear to you that the solutions are similar. So in this example, we have, like, a retrieval task and we have, like, an aggregation task, and they’re very different query. Like, the domain is just completely different. And when you train a regular language model over these two tasks, the trajectories look very different. And so what you’re relying when you. Like, let’s say you use Pi or Claude Code or something, which is not in this blog, but we do have these results. You’ll find that, like, these harnesses distinguish too much between these problems, even though the solutions are the same. And so one thing that we find when training RLMs is that, like. When it sees these problems, it’s the same. And the reason it sees these problems as the same is the sub-agent sees different problems, but the sub-agent is solving an easier sub-task, and so you’re confident it’s smart enough to do it. But for the base overall strategy, they end up looking the same. And so when you train on the left task, for example, it can immediately solve the right task. And so if you go down to, like, the plots that we have, one thing you’ll kind of observe is that the. as you just naively train your RLM on these tasks, they naturally learn to generalize, for example, to longer tasks, because the strategy is basically the same. You’re just modifying, like, a length variable. And this actually also holds for tasks that are different, and they’re, it’s not even-- they’re not different across length. They’re completely different tasks, math tasks versus writing tasks. But the solution, the, like, meta high-level solution is the same. And so when you train the RLM on one of them, it generalizes this behavior to the second one. And there’s no magic here. I guess maybe that’s the thing that I wanna kind of stress. Like
Vibhu [00:36:38]: How would you kind of verbalize what they are learning? So I think in here you say you train it on short tasks, they generalize to stuff 8–30x longer.
Alex Zhang [00:36:47]: Yeah.
Vibhu [00:36:48]: They are learning how to solve these type of problems, or what’s the, what’s the core thing they’re actually learning?
Alex Zhang [00:36:52]: Yeah. They’re learning how to solve these types of problems at a certain length. And it turns out that when you take the strategy that they learned, it is directly transferable to the longer length. Like, they’re
Vibhu [00:37:05]: Yeah.
Alex Zhang [00:37:05]: Effectively the same program. And this is something that is really exciting because what this kind of implies is that when you have a corpus of data or environments that you train your model on, the hope is that, like, A, you can train on less environments and generalize to more than what existing models can do through or harnesses through existing kind of, like, naive training. But B, also, you still wanna use all the data you have. So when you train on these tasks, like, hopefully it generalizes to a wider class of problems. And why this is also even more exciting, at least in the context of RLMs or any recursively calling system, is this argument holds inductively. So, like, I’ll give you an example, because I said, I claimed that competitive programming and GPU optimization use very similar skill sets. The model can. the harness potentially, you might have to nudge it a certain way, but it can learn that, like, “Okay, how I’m gonna go about solving this GPU programming task is very similar to what I learned for competitive programming. So I’m gonna list out a set of solutions. I’ll, I’ll, like, spawn subagents to list out promising solutions, and then I’ll, like, write this loop to go through and check these solutions, maybe evolve them, and, like, evolve them against a verifier.” And between these two tasks, this looks the same. But what the subagents are doing are maybe, like, unique and something, like, different. But even what the subagents are solving might actually also be of the same form, right? Because it’s like a, it’s a recursive argument. And so what I’m trying to get at with this whole blog post is just that, like, we should rethink what the role of the harness is, because harnesses can actually greatly increase the generalization capability of your model and the amount of data that it’s given. And this is not exclusive to RLMs. I think there is a wide class of harnesses that are yet to be discovered that actually can yield similar properties. And to extend this argument a little bit further, I think what you can also extrapolate from this is, like, if I look at an RLM, what are the components of an RLM that are actually necessary, and can I actually just directly train a model to do this? Can I train a model to act as an RLM implicitly in its forward pass? It’s a really weird thing to think about because, like, you might say like, “Oh, code is non-differentiable, blah.” But there are many approximations of this behavior that we will start to uncover. And I think, like, we will see beyond just, like, I’m gonna design a new coding harness that uses a special form of compaction or something. I think we can be a lot more creative here. Like, there’s, there’s so much we can do with these language models that I think we are just not doing. And I’m, like, very excited about this because I think, like, I think we can get serious gains from very opinionated and good harness design that lends itself better to scale. And what by this is, like, the RLM, for example, is a very primitive inductive bias. Like, there’s nothing super special about the design other than the fact that it’s very different than what we currently do. But this may potentially scale much better with, like, the data and the environments that we have available to us.
Vibhu [00:40:20]: I guess the, opposite thing that people would probably ask is current harnesses Are very generalized towards coding, which people see works for a lot of domains. Cloud code is being used for design, presentations
Alex Zhang [00:40:33]: Yep.
Vibhu [00:40:33]: Everything. MuseSpark, Grokbot.
Vibhu [00:40:36]: These are very simple, non-opinionated harnesses that are good at code, and that is also scaling out. what’s the example of how we improve those, I guess?
Alex Zhang [00:40:48]: Yeah. let me bring up another paper, which came out very recently. It’s like the harness tax paper. I think it’s by Arena. I really like this paper because it puts forward a prior that I had, which is basically that, like, most
Vibhu [00:41:07]: So it confirms the prior.
Alex Zhang [00:41:08]: Yeah. Like, most harness choices don’t matter because
Vibhu [00:41:12]: Yeah.
Alex Zhang [00:41:13]: All of these harnesses are the same. But I will say, like, Grokbot, for example, is actually quite different, I think, from my understanding, than how some of these other harnesses have been designed, and I like that a lot. and I think it’s clear from here at least that, like, I’m pretty sure. Anthropic or OpenAI are exclusively training on their harnesses. They’re probably not training on their competitor’s harness. I’d assume not, because I don’t know why they would do that. But
Vibhu [00:41:38]: But, this is a thing you see in open models, right? Like, Qwen is really good at using open code.
Alex Zhang [00:41:43]: Yes.
Vibhu [00:41:43]: They need to train in harnesses. Old Gemmas were notoriously bad at this.
Alex Zhang [00:41:47]: Yeah.
Vibhu [00:41:48]: Models are good, but you need to train in a harness.
Alex Zhang [00:41:50]: I think, though, as models get smarter, or, like, as they get better, this distinction becomes, not that important in the sense that, like, if you take Astra and you put it inside of open code, like, it’s not gonna go crazy, because I think it’s, like, just sufficiently good. And so I say this because the only benefits between these different harnesses is just cost, for the most part. And I think, like, what I’m getting at, Sugru, is if you plug these models into RLMs, though, they’re not that good still. They’re okay. And I think it’s mainly because the types, like the class of harness that we are training around is this class of harness, this, like, pi loop, this, like. I like to call it trajectory as a prompt, which just means, like, you keep the whole trajectory of the rollout as the context that your main model is using. Even if you use subagents, it’s still, like, kind of this form. And I think we’re going to. If we want to explore new harnesses, like, there needs to be teams that are dedicated to actually running meaningful experiments over, like, scaling out new harnesses, like maybe post-train scaling out on different harness designs. Like, I think we can actually get very meaningful knowledge or gains from doing this kind of thing, whether it’s an RLM or whether it’s something different. And that’s exciting, ‘cause I think, for example, if you train a lot on. Fable for a long time was the best model for RLMs because they had dynamic workflows, and it was pretty obvious that, like, this was a capability that was somewhat trained in. Even if the model was still, like, a little dumb, like, in the RLM harness, it still worked a lot better than other models did. Astra is now also, like, good enough at doing these things. But
Swyx [00:43:33]: Wait, is this where we see that Fable is the best for RLMs, or is there some other
Alex Zhang [00:43:38]: Oh, no, these are all internal results, I guess.
Alex Zhang [00:43:40]: Yeah, I don’t, I don’t have them
Swyx [00:43:41]: Okay
Alex Zhang [00:43:42]: Public right now. But in general, like, I think you can, You can very easily tell that we have not optimized for RLM, like, workflows yet. And I think if we get models that do this correctly, they will be a lot more efficient as well. you can kind of just think through, like, why this is the case, right?
Vibhu [00:44:03]: I think this is the point where you have to give the ten-second what are RLMs.
What Is an RLM?
Alex Zhang [00:44:07]: Oh, yes.
Vibhu [00:44:07]: Because there’s a lot of listeners here that
Swyx [00:44:09]: Yes. We’re, we’re assuming a lot of knowledge.
Alex Zhang [00:44:10]: Yes.
Swyx [00:44:11]: Also, I think you. But you have set some context
Vibhu [00:44:13]: Yes
Swyx [00:44:13]: So you can. Like, with everything we just said
Vibhu [00:44:15]: Yes
Swyx [00:44:15]: Can we have a clean, crisp definition of RLMs?
Alex Zhang [00:44:18]: Yes. Okay. I want to go back to the blog, the
Vibhu [00:44:21]: Yes
Alex Zhang [00:44:21]: The compositional generalizers blog. This one. Okay.
Swyx [00:44:24]: Okay.
Alex Zhang [00:44:24]: This is, like, the best. I wish I had this in the original paper. An RLM is basically just a harness design where the only tool in the harness is code, which is this programmatic subagent calling thing, where it has the option to call itself as a tool, and it has other tools. But all of these things are functions in code, and the context that it’s dealing with is always stored in some memory inside of this code environment. So this could be a file system. Like, this could be, like. I’ll give you an example, Prime Agent. The trajectory of Prime Agent, like the context, even the. when you compact and do all these things, is stored on disk. So the model can always reference its original context, even if it’s compacted, and all of its tools are run inside of, let’s say, like, a Python REPL or a Bash REPL. And so this. it’s like this very primitive abstraction. And I would say, like, where most harnesses differ is, A, context offloading is not done that, like that. if you look at Prime Agent, by the way, Prime Agent does context offloading, but not all the way. So, like, it still maintains the standard cod code, Codex loop of, like, trajectory as a prompt where you compact, but it has the additional kind of, like, the context is offloaded, and it only has. The unique point of Prime Agent is that the only tool is IPython. So this is, like, the very kind of generic abstraction around, like, RLMs. Yeah.
Vibhu [00:45:59]: Concretely, what’s the core thing RLMs are trying to solve?
Long Context, Composition, and Locally In-Distribution Tasks
Alex Zhang [00:46:02]: Yeah.
Vibhu [00:46:02]: At one point, I think when it first came out, it was context.
Vibhu [00:46:06]: I’ll pass the question.
Alex Zhang [00:46:06]: Yeah. So the original problem was harnesses have a really bad time dealing with long context. They typically were used to only really do them for, like, specific things like code. Like, they could deal with your code base because it was trained on it. But now it’s more around what this blog is talking about, which is Compositionality and the fact that, like, I think harnesses. We want to have language model systems that have much more control over the actions they make at every step. And what by this is tool calls are very limited because you have to invoke them every turn. Like, you have to invoke tool A, then tool B, then tool C, and there’s no central context that you can kind of draw back from. And RLMs are specifically, like, designed around composition and having, like, a central context that you can always draw from. and this context is, like, designed around the existing language models. Another, like, very similar example actually in design is, like, agent swarms, for example, the Hugging Face incident. Like, these agent swarms have, like, a message board that they learn to communicate over. And this message board, in some sense, is the shared context that they, like, act over. And RLMs basically say that, like, the best way to communicate through this is in code. Like, you write the code to do this, and it’s, it’s because these models are so good at writing code. like, we wanna take advantage of that fact.
Swyx [00:47:37]: And then for this compositional thing, up to and including generating your own harness, specific for the task.
Alex Zhang [00:47:44]: Exactly.
Swyx [00:47:45]: Right?
Alex Zhang [00:47:46]: I think we will start to see that if you go up to, this figure. Okay. So we talk about this idea of locally in-distribution tasks for a harness, and it is like a, an idea on top of, like, in-distribution tasks. When we think about language models, an in-distribution task is just a task where, like, the prompt is something that the model has either seen before or, like, has seen some version of it. Most harnesses work like the one on the right, which is they keep appending the trajectory as a prompt. And so eventually, unless you’re Anthropic or OpenAI and you train on, like, kind of these, like, user trajectories, most of these things end up being out of distribution for the most part. But locally in distribution is basically the compositional argument of if an RLM breaks down its computation into, like, kind of a meta-harness of sorts or, like, a program that involves subagents that, like, look at a local problem, every individual language model call over the course of this task is in distribution, even if the entire task itself is out of distribution. And this is a very desirable property, I think, for obvious reasons. Like, if every task is in distribution for each individual language model call, you will probably get to the right answer. so.
Swyx [00:49:00]: The logical limit of RLMs is RLLMs where, like, you not just, you don’t just write the harness, you also train a custom model for
Training RLMs and Smarter Harnesses
Alex Zhang [00:49:11]: Exactly.
Swyx [00:49:11]: You collect data, everything.
Alex Zhang [00:49:14]: Yeah.
Swyx [00:49:14]: Like, it’s a fully automated AI researcher inside of your harness.
Alex Zhang [00:49:16]: Yeah. We will see where the training of RLMs goes. I will say, as an academic, I am not working on this at MIT, or at least in the scaled sense, because I can’t afford to. but there are companies out there that are working on this. I think Prime and Select is very clearly working on this, and it’s very cool. Like, I’m, I’m very excited to see. Maybe we’ll observe, I don’t know, but maybe we’ll observe better, like, post-training scaling laws with when you train around a smart harness. Maybe we’ll even see smarter harnesses that come out and, like, they work better around these kinds of principles.
Swyx [00:49:50]: What is a smarter harness? Like, that doesn’t mean anything. You just said they’re all the same.
Alex Zhang [00:49:54]: No. What, more of what is, like, Claude Code, Codex, Pi, et cetera, are all the same in that, like, when you break down the logic of the harness, it’s, like, virtually the same thing.
Swyx [00:50:06]: Yeah, two calls in a loop or
Alex Zhang [00:50:07]: Yeah
Swyx [00:50:08]: Whatever.
Alex Zhang [00:50:08]: But With RLMs and with other harness abstractions, it looks very different. And this is where I think you really distinguish. it’s, it’s in the same way that, like, I think with language model architecture choices, a lot of architecture choices end up kind of looking the same when you, like, scale it out or, like, it. The differences end up being, like, somewhat minor in terms of. for a lab it’s not minor, but, maybe one model converges better than the other one, like, slightly. But in general, like, if you were. So for example, pre-training scaling laws only hold because the architecture choices we have are somewhat stable right now. But if you were to completely change the architecture, pre-training scaling laws probably don’t hold. Or, like, these kinds of. This, like, power law is gonna look very different. And it’s like the same thing with harnesses. Like, I think all the harnesses we have right now, for the most part, roughly look the same, but there are some exceptions to this, I think, that are coming out.
Swyx [00:51:06]: I was gonna say, I actually, one of the things that I’ve been more interested by, like, talking about PhD students who take big risks, is that people have been. People also pursuing the other side, which is pre-training scaling laws don’t hold if you change data. right now it’s just raw, unstructured text, corpus of internet. What if you had a better data representation to train on? Then yes, your scaling law would change as well.
Alex Zhang [00:51:28]: Yeah.
Swyx [00:51:28]: So there’s architecture, there’s data, and, whatever else, you can think about. Well, so I just wanna get back to this. it all makes sense. It’s, it’s very interesting how you sort of recurse up and down the stack from, like, very conceptual to, like, not like, well, this is where we are today.
Prime Agent and Opinionated Harness Design
Swyx [00:51:43]: But, like, yeah, obviously, it can scale up and down. I guess, I’m curious, how did you start working with Prime? Is Prime taking on more work with this? Is this their answer to Hermes agents?
Alex Zhang [00:51:56]: Mm.
Swyx [00:51:56]: You mentioned Grokbot is a little bit different. I just wanted to, like, namecheck all these guys and get your thoughts on each.
Alex Zhang [00:52:01]: Yeah. I got involved with Prime, after they released a blog post, by the way, not affiliated with me at all, about, like, how they believed RLMs were kind of the future. And I had a friend that was working there, GPU Mode, Matei. Like, we got in touch, and I think I agreed with a lot of the researchers there and, like, what they believed about harness design. Like, I was very impressed, I think, that, like, they understood the purpose of the RLM paper, which is not necessarily just to say that, like, we’re solving long context tasks, but actually, like, we want more opinionated harness designs.
Swyx [00:52:39]: Yeah. There’s always, like, the result of the paper that you choose to highlight
Alex Zhang [00:52:42]: Yes
Swyx [00:52:43]: Versus the actual point.
Alex Zhang [00:52:44]: Yes. As I would love to talk about, like, the incentives of academia and, like, the things around, like, why it’s kind of flawed and all the issues, and we’ll get back to that. Yeah. So anyways, I love the guys at Prime. So we kind of had been. After we decided to work together, we decided to look into training in RLM and also build this kind of RLM harness and kinda see where we can take it. That is how, like, Prime Agent came about, and I think the reception for Prime Agent has been pretty good. Like, the one thing I was worried about with Prime Agent is that none of them, at least at the time when we were building it, none of the models were that good at doing RLM stuff. So this was, like, pre-Fable, pre-Astra.
Swyx [00:53:27]: I guess, I think to take a step back, can you explain what Prime Agent is, how it’s different than
Alex Zhang [00:53:32]: Yeah
Swyx [00:53:33]: A traditional, Claude Code, what people would expect harness?
Alex Zhang [00:53:36]: Yes. So Prime Agent, I think I mentioned this a little bit earlier
Swyx [00:53:40]: Yeah, there was the diagram. Yeah
Alex Zhang [00:53:41]: Is basically. it is a. A harness on top of Pi, like Pi Mono, which is-- Pi Mono, for context, is like the, like a
Swyx [00:53:50]: Core agent
Alex Zhang [00:53:51]: A minimalist
Swyx [00:53:51]: Yeah.
Alex Zhang [00:53:52]: Yeah, like harness. I use Pi as the reference for everything because I think all other harnesses are basically just Pi.
Alex Zhang [00:53:57]: But it is Pi, except we explicitly restrict IPython to be the only tool that’s available to it. Every other tool gets loaded in as, like, a Python module, or like a Bash kind of script that it can run. So it uses the core RLM abstraction on top of Pi, and then it also has this continual harness thing, which is, Seth, he’s another PhD student. This is a thing that he used to get language model harnesses to play games. Like, so he worked a lot with Joel, who is the, like, Gemini plays Pokemon guy. And continual harness is also, by the way, very simple. I quite like it. It basically is this design, principle around, like, what parts of the harness can you let the harness itself modify? There are certain pieces that, like, you’ll let it modify its own skills, the subagents available to it, what the system prompt to the model is. And continual harness is available basically as a tool inside of the IPython kernel. And so that’s what Prime Agent is, like, how, what it’s designed around. Everything else in Prime Agent is like
Swyx [00:55:05]: Standard.
Alex Zhang [00:55:06]: Standard.
Swyx [00:55:06]: Standard.
Alex Zhang [00:55:06]: Right? Yeah. I think what is, what I really liked about it, and we got kind of lucky, is that, like, a lot of the new frontier models actually work really well inside. And actually, even a lot of the open-source models work really well, at least some of the newer ones. And there is another thing in Prime Agent I should highlight, which is that, like, we have a very particular agent-to-agent communication system or, like, framework, which is because RLMs tend to spawn many subagents, we want a way for subagents to communicate with maybe the root or with each other. And so there are some design decisions around, like, what each subagent is allowed to talk to, how it does it. Again, everything is in code, so it writes the code to do this kind of communication, which I think is really cool. And then there’s, I guess, persistent subagents is another thing that was kind of added, which is the subagents, they can last beyond, like, the standard runtime of the actual, like, original agent. And you can go into that subagent, you can prompt it more, like, you have more visibility and flexibility into what is kind of going on.
Swyx [00:56:10]: This is my number one pain with Codex right now. They, their subagents are just very ephemeral And they actively discourage you from using it for long-running things.
Alex Zhang [00:56:17]: Yeah. Yeah. Which I think it makes sense. Yeah. I
Swyx [00:56:21]: So the trick is just externalize to a file system, right?
Alex Zhang [00:56:24]: Yes. Yeah. That’s
Swyx [00:56:25]: Like, that’s the trick.
Alex Zhang [00:56:25]: That is the big trick.
Alex Zhang [00:56:27]: Yeah.
Swyx [00:56:27]: And well, and also, like, force everything to run through code. trust the model
Alex Zhang [00:56:30]: Yep
Swyx [00:56:30]: That can write code, and it’s gonna write its own harnesses itself. So is Prime gonna take on, like, training, post-training custom models for this? Is this a one-off collaboration between you guys, that’s it? Like, what’s
Alex Zhang [00:56:41]: Yeah. They are training a model, intern-- I think they were pretty public about this actually
Swyx [00:56:46]: Yeah
Alex Zhang [00:56:46]: Back in March.
Swyx [00:56:47]: Clearly it is their business.
Alex Zhang [00:56:49]: Yeah.
Swyx [00:56:49]: Yeah, so.
Alex Zhang [00:56:50]: Yeah. they’re, they’re showing that they can train it on their kind of hosted training stack. But no, so for model training, I’m, I’m not involved with them on that. The main reason is just I have other things in the PhD I wanna work on. I think, like, there are many other big bets to take,
Swyx [00:57:03]: Ooh
Alex Zhang [00:57:03]: Outside of just RLMs. some
Swyx [00:57:06]: Ooh
Alex Zhang [00:57:06]: Some I don’t know how much I can share yet. but in general, like, I think, I actually think one of the luxuries of being a PhD student, genuinely, is that there’s so many big bets to take. most of them will probably yield nothing, but it’s a really exciting time to be in research, especially because I think most progress in the field has been a little bit boring. Like, I’m not saying the outcome is boring, but the process of doing these things tends to be quite boring. and so there is kind of this question of, like, what do we wanna do next? but
Swyx [00:57:39]: Yeah
Alex Zhang [00:57:39]: We can talk about that later.
Third-Party RLM Work: Harvey, Headlong, DSPy, and ARC-AGI-3
Swyx [00:57:41]: Yeah.
Alex Zhang [00:57:41]: Yeah.
Swyx [00:57:41]: Okay, I wanna close out a little bit more of your research, and then we can,
Alex Zhang [00:57:44]: Cool. Yep
Swyx [00:57:44]: Start putting it out. since you released RLM, a lot of excitement about it. Any secondary third-party work that you wanna shout out as, like, that you guys should take a look at this?
Alex Zhang [00:57:53]: Oh, yeah. So, Harvey, the legal AI company, released a blog post, not affiliated at all, but they post-trained an RLM on their, like, legal work, which often involves a lot of, like, sifting through documents and kind of looking through, like, a variety of, specific information that maybe is not so easy to retrieve with, like, a pure retrieval system. And they show, like, really good results. It’s very exciting. I was shocked that they worked on this. They did not tell me, so when this came out, I was like, “Oh, that’s awesome.” So there’s this one I think is super cool and what they’re doing there. I think this is a collaboration with Base 10, by the way, as well.
Swyx [00:58:33]: Yes, this was Base 10.
Alex Zhang [00:58:34]: Headlong, which is law, the Law Institute’s kind of. it is their, like, persistently running harness. it’s very cool that
Swyx [00:58:42]: Oh, they renamed it? They used to call it something else.
Alex Zhang [00:58:45]: It was like Auto
Swyx [00:58:46]: Terminus.
Alex Zhang [00:58:46]: Yeah, I know. They’ve gone through. Yeah.
Swyx [00:58:49]: All right.
Alex Zhang [00:58:49]: So this is Andy Konwinski’s big project. it’s super cool. I love Andy. I don’t want to downplay what they’re doing because they’re using the RLM abstraction, but they’re doing something much cooler than the RLM, which is like they have a system that kind of what they call, like, thinks persistently. So even when you don’t query it has a way to, think through problems that it has in its context.
Swyx [00:59:16]: Oh, so it’s just like a always-on type thing.
Alex Zhang [00:59:18]: It’s like an always-on thing, but it’s, like, not that expensive. Like, they control the token costs, to make sure it’s not, like, burning through all your credits. This is super cool. I’m trying to think. There are many. Actually, if you go to the RLM, GitHub page, there’s a bunch of things I’ve linked, below. There’s a ton of really cool kind of things that people have been doing. Axe is another really cool one that I think it’s just by this one guy. It’s like a harness around DSPy and RLMs. DSPy also has an RLM. Oh, the last thing I’ll shout out is on ARC-AGI-3, I believe, there were a lot harnesses on their, like, Kaggle competition, like the official one, not the, like, public primates, like one that, or like what people have evaluated on. They all, like, claim to use or they reference, like TUFA, for example, some form or some inspired form of the RLN abstraction in their harness, which is really cool. I think it’s, This is where-- this is exactly the setting where you would see a lot of benefits from composition and using code and combining, like, neuro symbolic systems with AI. And so
Swyx [01:00:27]: Yeah.
Alex Zhang [01:00:27]: Yeah. Very cool.
Swyx [01:00:28]: We love a good neuro symbolic reference.
Agent Swarms, Unsolved Math, and What the User Should See
Alex Zhang [01:00:30]: Yeah.
Swyx [01:00:30]: You also, mentioning ARC-AGI-3, OpenAI comes out and says, “We’re at 99.9% on this.”
Alex Zhang [01:00:36]: Yep.
Swyx [01:00:37]: They also say, “We solved Navier–Stokes. We just threw a model at it.”
Swyx [01:00:40]: There’s some debate around whether or not it’s just model.
Swyx [01:00:44]: Are they using an RLM? Do?
Alex Zhang [01:00:46]: I would guess probably not, unless you say, like. I’ll, I’ll be, I’ll be careful here because, people debate what is an RLM, what is not an RLM. It’s somewhat clear that what they used is some kind of swarm of agents with a shared, some shared context, like some shared file system. And, like, this is very much in the spirit of RLM stuff, but I think there’s a, there’s a lot of, like, more clever things that they did that’s not maybe related to the RLM itself. I agree a little bit with the idea that, like, a harness is not that necessary for what they did. The way that I would put this is that I think a model, like a GPT-6 Astra type thing, is technically smart enough, conditioned on the right information, to come up with a proof for these very difficult problems. Now, how you get to that information is a giant question mark. And in their case, it probably came down to, like, a very long search over, like, many of these sub-age-- or many of these, like, agents in the swarm and maybe also, like, researchers cond-- I’m, I’m actually not sure about this part, but putting in, like, their kind of intuition as to, like, what you should explore and things like this. And ultimately, like, this produced some information that some agent was able to take to finish the proof. And so in that sense, like, I think, was the harness that important? No. And I think what this is pointing at is, like, the specific details of a harness do not really matter, and I think that’s also what that, what the harness task paper is pointing at, which is that, like, beyond the user’s feeling of the harness, realistically all that matters is just, like, how are you composing these agents in a meaningful way to get to the final answer? And maybe that’s what, like, swarms and all these things are really about. And so from my POV at least, if we start to think about, like, for user use cases, what do we want out of harnesses and things like this? Like,
Alex Zhang [01:02:51]: We want to take the good parts out of these, like, the Claude codes, the Codexes, like the stream that people like to see. But, like, under the hood, whatever is running can be some really weird, complicated swarm of agents that, like, ultimately come up with an answer. The user doesn’t wanna see that, though, obviously, right? Like, it’s, it’s not legible information. And so I-- this was another kind of thing in the spirit of RLMs, like recursive language model. It sounds like it’s a language model and, but it’s not a language model architecture. But the reason for this is, like, I think we will start to see in the future probably one day, and I wrote a blog about this, what we think of as a language model, like the thing that we query, might actually be like a swarm or like a scaffold or like some Weird harness design that scales very well, but the user just doesn’t see it. Ultimately, all the user sees is some front-end version of this harness. And yeah, I think it’s a relatively safe bet at least to make that this is what we will see.
Swyx [01:03:48]: Yeah.
Alex Zhang [01:03:49]: And this maybe goes back to the limitations of the base transformer. Like, obviously if you just took a base transformer and you said, like, “Solve Navier–Stokes,” or something, it’s not gonna do it. Like, yeah, we all know this is not what’s gonna happen. But yeah, I think this is maybe the more interesting part. and maybe the claims around, like, did the harness matter is more around this, of, like, just arbitrarily pointing models, like, or agents at a growing kind of context of information maybe is just enough to solve very difficult problems. that I can buy.
Swyx [01:04:20]: While there’s a lot in there, I do wanna say, the amount that OpenAI spent is semi-public. It’s, 10,000 agents in 88 hours.
Alex Zhang [01:04:29]: Oh, yeah.
Swyx [01:04:29]: 130 billion output tokens, which is estimated to be about 40 million dollars in public pricing.
Alex Zhang [01:04:34]: Surprisingly, actually, like, less than I thought.
Swyx [01:04:37]: Yeah, not that much.
Alex Zhang [01:04:38]: Yeah. Yeah.
Vibhu [01:04:38]: 130 billion output for the final, but as you said, there was a lot of context being passed around.
Vibhu [01:04:44]: It’s more than double that in just the total agent messages being sent.
Alex Zhang [01:04:47]: Yeah.
Swyx [01:04:48]: Yeah.
Alex Zhang [01:04:48]: Yeah. Yeah.
Swyx [01:04:49]: I think, one thing I was honored to bring up also was Cursor as far as, like, swarm stuff is concerned.
Swarm Architectures, Coordination, and Token Efficiency
Swyx [01:04:53]: This is slightly older, meaning February, which is ancient.
Alex Zhang [01:04:57]: Whoa.
Swyx [01:04:57]: But if you scroll down all the way to the final sort of multi-agent architecture that they arrived at, it was basically an org chart of a normal software team. One thing I’m thinking about, because I basically. There’s, like, this, gather all function That you have to do with subagents or. It’s very similar to GPU programming, actually.
Alex Zhang [01:05:16]: Yeah. Yeah. Yeah.
Swyx [01:05:17]: And so, that’s a bottleneck. This is a bottleneck.
Swyx [01:05:21]: If there’s one main guy, that’s coordinating all the sub guys, then they have to, like, gather again and then re-coordinate.
Swyx [01:05:29]: That’s slow. That’s, that’s crappy. what a true swarm should be is everyone is just their own person.
Alex Zhang [01:05:35]: Yeah. Yeah.
Swyx [01:05:36]: Right?
Alex Zhang [01:05:36]: Well, I agree with this, and I think that there is a question to be had, though. Let me give an analogy, which is like, when would you use compaction and when would you use an RLM? And there are many settings where an RLM can solve maybe a more difficult task than compaction can, but you would prefer to use compaction in most cases because it’s cheaper and it’s quicker. And I think in the context of agent swarms, there is a similar thing going on of, like, I’m. Fairly certain that, like, 95% of the swarm is entirely useless, or, like, what it’s exploring is entirely u-- You’re just burning tokens. Versus in this setup, maybe not so much. I’m not sure. maybe it’s also the case here.
Swyx [01:06:17]: Everyone has a job. This is your board.
Alex Zhang [01:06:19]: Yeah.
Vibhu [01:06:19]: I think at some level that’s how problems are framed, right?
Alex Zhang [01:06:23]: Yes.
Vibhu [01:06:23]: So, like, if you have a search problem and you’re spanning out a bunch of subagents to do search, there’s gonna be a lot of useless information, right? There is one retrieval answer That you’re getting and you’re spanning off, but that is consciously understood, right?
Alex Zhang [01:06:36]: Yes. But there is kind of this question of, like, what is appropriate to solve for which problems? Like, what design-- in theory, OpenAI can use-- can package up this API and they’ll call it swarms, and then they’ll give it to you and they’ll be like, “Point this at any problem and we’ll give you a solution.” But maybe
Swyx [01:06:53]: Yeah, it’s called, it’s called pro, right?
Alex Zhang [01:06:54]: Yeah. maybe you’ll have to pay like 40 million dollars to get a result.
Alex Zhang [01:06:57]: And it’s like, well-- But it’s exciting. I will say, like, it is-- it’s very exciting that we even have the option to point 40 million dollars at a problem and solve it.
Vibhu [01:07:07]: Yes.
Alex Zhang [01:07:08]: But there is still kind of, a lot of research to be done in this area around, like, what is necessary. Like, what do we want to do? What design do we want? we probably don’t want everything to be a swarm, but, like, where do we draw the line? Like, can the agent design that or decide that for itself? Et cetera. So.
Open-Endedness and Research Without a Fixed Objective
Swyx [01:07:26]: Yeah. and then, just a quick check in case you have opinions on this. Have you looked much into open-endedness as a general category of problems?
Alex Zhang [01:07:35]: Oh
Swyx [01:07:35]: Meaning no prompts, just go.
Alex Zhang [01:07:38]: A little bit. so I was at Sakana for a summer, right after, or I guess right before my PhD, and that’s something that they work on a lot there. And I think there’s a lot of people even at, like, Recursive Super Intel-- There’s many of them now.
Swyx [01:07:53]: Yes, we just had Richard Socher on.
Alex Zhang [01:07:55]: Oh, yes. Yeah. So, like, Richard’s company and then also. Actually, wait, that might be the s-- it might be the same company. I don’t remember. Is Tim Lautenschlager also
Swyx [01:08:03]: Yeah.
Alex Zhang [01:08:03]: Okay. Yes, that company.
Swyx [01:08:04]: He’s the main co-founder. He used to be head of open-endedness for Google.
Alex Zhang [01:08:07]: Yeah.
Swyx [01:08:07]: Yeah.
Alex Zhang [01:08:08]: So I think with open-endedness problems, like, I view them as somewhat similar to even, like, unsolved math problems. Maybe that’s a weird way of framing it, but, like, I think a lot of the techniques in terms of, like, how people like, approach them are kind of the same. Like, evolutionary search is, like, very similar to just launching swarms of agents and hoping that, like, they come up with, like, an interesting s-- And this is what, like, AlphaEvolve and some of these other works did, like a year or two ago. But I think what maybe is not, And I’m not sure if this is what you were alluding to, but I think what’s not super clear in open-endedness style search problems is do we frame all of them as basically like an unsolved, very difficult problem, or, like, where the objective is clear? If that’s not the case, I still don’t know yet entirely what the value of it is. maybe you have other opinions. Like, I don’t have too many opinions on this, but at least from my time at Sakana, like, I got the sense that, like, we ultimately still kind of wanna approach things the way that, like, say, OpenAI approached Navier–Stokes. We want. We-- There’s still a lot of nudging in certain directions that we want to have to, like, get to the point where we have something interesting.
Swyx [01:09:29]: My version of it is, like, maybe it’s a split between basic science and applied science. Basic science, you’re researching for researching’s sake.
Swyx [01:09:36]: You just wanna understand things better. I have no idea if, like, there will be any application at all whatsoever, but, that, And then applied, you have a goal.
Swyx [01:09:45]: You’re, you’re trying to minimize loss in some way? and so, what I really, think, in terms of, like, the big bets that people have, what if there was no prompt? Like.
Alex Zhang [01:09:57]: I see.
Vibhu [01:09:58]: You just pick domain and let it
Swyx [01:09:59]: Like, you just, like, you just spawned in this, like, swarm of things and you’re like, “Hey, what’s up, guys? Like, what you guys working on?”
Alex Zhang [01:10:03]: Yeah.
Swyx [01:10:03]: And, like, you just decide
Alex Zhang [01:10:05]: I see
Swyx [01:10:05]: Like, this is an interesting problem.
Alex Zhang [01:10:06]: The biggest issue that, like. And maybe there’s a clear path to this, but when I was there at least, the biggest issue that people had with open-endedness was like, how do you ultimately choose at the end? How do you pick out the interesting stuff? Because when there is no. Like, maybe the agent comes up with a goal, but in a lot of cases, like, what they. And they have something called Fugu, I think, which is like a It’s like a model router type thing that was, like, inspired at least by this idea of, like, let’s pick a problem where maybe we can pick out the best solutions to something. In this case, it’s like pick the best model for this problem.
Swyx [01:10:41]: You’re the first person to connect model routing to open-endedness.
Alex Zhang [01:10:43]: No, yeah. But, so I bring this up because I think, like, with open-endedness, like, just generally the issue is, like, when we have this giant corpus of, like, slop, like, how do we sift through
Swyx [01:10:57]: Yes
Alex Zhang [01:10:57]: And find, like, the hidden gems? And, like, the solution to open-endedness really just letting models run forever and, like, finding. Like, just doing data gen-- just doing super high throughput data generation and then, like, asking agents to go through and, like, find meaningful things. Like, I’m not sure. Maybe that’s sufficient? Like, that would be, that would be cool.
Swyx [01:11:21]: To me, it’s, like, very interesting as a counter to basically all of machine learning Where you have a goal, to have no goal.
Alex Zhang [01:11:28]: Yeah. Yeah.
Swyx [01:11:29]: But, or, like, an ill-defined goal that you’re like, “Well, what about this goal?” And you’re like, “Well, okay, maybe.” And then you, like, sort of research more and you find
Alex Zhang [01:11:36]: Yeah
Swyx [01:11:36]: That is an interesting goal. ‘Cause, like, I think, like, finding the objective function, like you said, like, Jeff found an objective function That was interesting that no one was exploring.
Alex Zhang [01:11:43]: Yep. Yeah.
Swyx [01:11:44]: I think that is, like, similar to your message about grad students as well. Like, you stay in school because you are. you want to pursue open-endedness. If you want to, profit max and, like, join the, escape the permanent underclass, then you join a lab.
Swyx [01:11:59]: Right?
Alex Zhang [01:11:59]: Yeah. Yeah. It’s funny. I feel like I don’t, I don’t hear this discourse a lot. I’m, I’m in the East Coast, so it’s, like, a very different type of. But then when I. whenever I come here, it’s like, that’s always, like, the topic of discussion.
Swyx [01:12:12]: You cannot pay rent without doing this.
Alex Zhang [01:12:13]: Yeah.
Swyx [01:12:15]: Yeah, you’re getting priced out, guys.
Vibhu [01:12:16]: Yeah.
Swyx [01:12:17]: Okay. So, yeah, there’s, there’s all that. I don’t know if you wanna-- if it’s relevant, enough to talk about the mismanaged geniuses, which you were pulling up.
Vibhu [01:12:25]: No, it’s just on your blog. But I will poke on, Sakana.
Swyx [01:12:28]: Oh, Sakana? Oh, okay.
Vibhu [01:12:28]: Yeah. So they did in their, blog post, I talked to them about this as well. So one of the cool results of Ultra is they basically just let it loose on automated data science research
Sakana AI and Weird Research Bets
Vibhu [01:12:40]: With little to no human intervention. It’s just kind of making
Swyx [01:12:44]: Yeah, so this is auto research, which is a little bit more open-ended, and there’s degrees of open-endedness, and I agree with that.
Vibhu [01:12:51]: Yeah, separate than auto research with objective, this is just
Swyx [01:12:54]: Yeah
Vibhu [01:12:54]: Do stuff. But, it’s cool. They’re, they’re working on it for those that are interested.
Swyx [01:12:58]: While you’re bringing it up, actually, what is your take on Sakana? Like, what are they doing apart from being, the Japan one?
Alex Zhang [01:13:04]: Yeah.
Alex Zhang [01:13:05]: I actually love the people there. Like, I think they have a really smart team. and it makes sense. it branched off from, like, an earlier team at GDM, which was also kind of, I guess, doing this kind of, like, open-ended evolutionary research style stuff. What I liked about my experience there, at least, was that they did have that, like, mishap back, I forget, at this point when, but I think, like
Swyx [01:13:31]: You’re talking about AI scientists?
Alex Zhang [01:13:32]: No, the GPU kernel.
Swyx [01:13:34]: Oh, okay, yes.
Alex Zhang [01:13:35]: That one. Yeah.
Swyx [01:13:36]: People cannot forgive them for that. Yes.
Alex Zhang [01:13:37]: Yeah. And I guess, like, the AI scientists, like, there’s, there’s some criticisms of it that I don’t, I don’t work on that, so I have no kind of take on it. But I think in general, like, what I like about them at least is that they’re a little bit more of a researchy type lab. So, like, they don’t operate in the same space as, like, OpenAI or Anthropic. Like, for sure, like, definitely no. They do not. it’s pretty obvious probably that, at least when I was there, they do not have a big competitor model or something that, like, that everyone is using. But I think they kind of operate in some ways as, like, a PhD lab, which is cool. Like, and I think, like, David Ha is, like, he’s, he’s really smart. Like, I think he has a good sense of, like. Also, I think the market in Japan is also a little bit different for AI, and, like, who they’re targeting is slightly different than maybe what we’re used to here. But yeah, I like that they take kind of. A lot of their research is kind of weird, I think, when people view it? And I like that. Like, I think it’s
Swyx [01:14:34]: We should have more weirdness, yes.
Alex Zhang [01:14:35]: Exactly.
Swyx [01:14:36]: And you said different market. Just, is it, like, enterprise?
Alex Zhang [01:14:39]: Like, the way it works there is a bit different, like how deals happen and stuff like that.
Vibhu [01:14:44]: They do have a. I guess this page is originally in Japanese, but they do have a model
Alex Zhang [01:14:49]: Oh
Vibhu [01:14:49]: Specialized for the Japanese market.
Swyx [01:14:51]: So you didn’t know that.
Alex Zhang [01:14:51]: I didn’t know that.
Vibhu [01:14:52]: I didn’t know it too.
Alex Zhang [01:14:53]: Did not know this.
Vibhu [01:14:53]: I also have personal friends that know the team.
Alex Zhang [01:14:56]: Yeah.
Vibhu [01:14:56]: So there is a. Even from the sense of a way that you speak culturally Responses are tuned towards that. This is not like it’s frontier on benchmarks. It is a cultural appropriate model for them, and then they have, like, chat and all that.
Alex Zhang [01:15:11]: Yeah.
Vibhu [01:15:12]: But, to mirror your point, there’s also, like, how should education look like? And someone wants to work on it, and they’re a very PhD lab of, “Do your thing. Why not? We have money. Go research.”
Swyx [01:15:22]: Oh, they say it’s a Kimi fine-tune. That’s nice.
Vibhu [01:15:24]: Oh, there you go. Kimi.
Swyx [01:15:25]: Good for them.
Swyx [01:15:26]: Yeah, and, speaking of Kimi, right, like another, just a grad student that spit out and like, yeah, I just have this, like, Kimi delta attention that wants-- that I wanna work on.
Vibhu [01:15:36]: Yeah. Yeah.
Swyx [01:15:36]: And, like, somehow managed to make Moonshot. Don’t understand it still.
Alex Zhang [01:15:41]: Yeah. Well, he’s, he’s super cracked, at least my understanding. I think in general, like, a lot of, a lot of the Chinese labs have done really cool work.
Kimi Swarms, Dynamic Workflows, and Convergence
Swyx [01:15:51]: Yeah.
Alex Zhang [01:15:51]: Like, yeah.
Vibhu [01:15:52]: Any thoughts on Kimi agent swarms?
Alex Zhang [01:15:55]: Yes. one thing I will say is whatever OpenAI is doing with their agent swarm is, like, clearly the right thing to do. You have to kind of think about it this way. Like, nothing, especially without, like, a very smart harness design, which I don’t, I don’t think anyone really has so far, Nothing comes so easily for free. For example, like, the agent swarm design is not something you can just take for granted. Like, it’s not like GPT-6 Astra is just super smart and then it just got agent swarms running well. They clearly trained. the Hugging Face incident was them training a system to be like a swarm. and I think, like, clearly they’ve done something really well to the point where you can throw 40 million dollars and solve an unsolved problem. And I think with the Kimi agent swarm thing, like, at least from when I read it just came off as like, this is interesting, but I don’t actually know whether or not this can solve anything novel.
Swyx [01:16:59]: Yeah, they just. They were like, “It makes spreadsheets for you.”
Alex Zhang [01:17:01]: They kind of were just like, yeah, like, here is a, here is a swarm that, like, kind of does stuff, and it’s cool. but. And I will say the same thing with dynamic workflows. I actually kind of think the dynamic workflows release was sort of a flop. I don’t know how you guys feel about it, but I think it’s like. my understanding is, like, it’s not used that often or, like
Swyx [01:17:21]: It’s just very expensive.
Alex Zhang [01:17:22]: It’s too expensive and like
Swyx [01:17:22]: It’s ultra code. It’s basically like, take over my bed.
Alex Zhang [01:17:26]: And it doesn’t I’ve tried it, and, like, it doesn’t act in the way that, like. Again, I’m not, I’m not the biggest OpenAI, like, stan or something, but I think whatever they did was very impressive. Like, they somehow managed to get a way for this swarm to actually act
Swyx [01:17:41]: I see
Alex Zhang [01:17:42]: Towards a goal. And yeah.
Swyx [01:17:44]: I see.
Alex Zhang [01:17:44]: It’s very difficult.
Swyx [01:17:45]: I see. So, like, efficiency of the multi-agent swarm is the objective function here.
Alex Zhang [01:17:51]: Yeah.
Swyx [01:17:51]: Right?
Alex Zhang [01:17:51]: Yeah.
Swyx [01:17:52]: Like, how much of this is slop? Like, this is a lot of slop.
Alex Zhang [01:17:55]: Yeah, I think so.
Swyx [01:17:55]: OpenAI is less slop.
Alex Zhang [01:17:56]: We take for granted what it means for a swarm to converge to an answer.
Swyx [01:18:00]: Yeah.
Alex Zhang [01:18:00]: It’s just like. It’s not something we take for granted.
Swyx [01:18:02]: Yeah. Yeah. We’ve, we’ve done one pod with Noam Brown and, like, his
Alex Zhang [01:18:06]: Oh, yes, I remember. Yeah
Swyx [01:18:06]: His thing. His whole thing was like, okay, like, we’ve worked on a lot of, like, competitive agents. we’re working on collaborative agents.
Alex Zhang [01:18:13]: Yeah.
Swyx [01:18:13]: And, like, that’s now called a swarm.
Gemini, GDM, and Harness Engineering at Scale
Alex Zhang [01:18:15]: Yeah.
Vibhu [01:18:15]: I would do a quick poke. Do you have any thoughts on Gemini? They were also IMO gold. Like, there was a time where they were getting agents to reason for a long time and.
Alex Zhang [01:18:26]: Yeah,
Vibhu [01:18:28]: Is it too old to think about?
Alex Zhang [01:18:29]: No. I think it’s a little bit blown out of proportion. Like, Gemini. For, again, I, so I should preface by saying I haven’t worked at any of these places, so take this with a grain of salt, right?
Vibhu [01:18:39]: You have strong opinions on agent harnesses
Alex Zhang [01:18:42]: Yeah
Vibhu [01:18:42]: And, this was
Alex Zhang [01:18:43]: But, let me just say this first. This work was really impressive. I think what they showed here was, like, they took a time when the models weren’t that good
Vibhu [01:18:53]: Yes
Alex Zhang [01:18:53]: And they managed to be very smart about, like, what the harness does. I remember for this at least, like, yeah, like AlphaGeometry, I guess that was a year before this, but it was very cool. They took it to the max, and they, like, designed, I don’t know. I’d say I’m not too big on competitive math, but I think, like, GDM, it’s sort of a shame. Like, everyone I’ve talked to about GDM kind of has the same opinion, which is that it’s way too, like, bureaucratic. Whatever is they have the talent and the resources to do almost anything, but, like, I don’t know, until they figure that part out, like. Nothing against anti-gravity, for example, but, like, I don’t know anybody that uses anti-gravity. And so I’ve tried it once, and it’s. I don’t see a reason to switch to it. and I think for whatever reason, like, they’ve been struggling with this, so, yeah.
Swyx [01:19:40]: Yeah. Well, a lot of people dogged on Meta for a long time until they started
Alex Zhang [01:19:44]: Yeah, and they recovered
Swyx [01:19:44]: Coming out. And, like, I think, Google’s going through that phase right now.
Alex Zhang [01:19:47]: Yeah.
Swyx [01:19:48]: And, it’s, it’s just you gotta stay alive and
Alex Zhang [01:19:51]: Yeah.
Swyx [01:19:52]: I wanna focus back on
Alex Zhang [01:19:53]: Yeah
Swyx [01:19:53]: Just, like, your thoughts, just general. we can talk about speculative PTC
Speculative PTC and Parallel Tool Execution
Swyx [01:19:58]: Mismanaged geniuses, or just, like, throw away all this and just talk about whatever else.
Alex Zhang [01:20:03]: Okay, let’s, let’s talk about mismanaged genius for a little bit.
Swyx [01:20:06]: Yeah.
Alex Zhang [01:20:06]: I, the only comment I’ll say on speculative PTC is that it’s a really simple idea. It’s almost, like, obvious that this should be done, and, like, there’s not much more to talk about it. Like, I think it’s just, like, you should just use it for, like, coding. Like, anything with programmatic agent calling, like RLMs or Kodak, like, yeah, it’s like, it’s like a no-brainer.
Vibhu [01:20:25]: What’s the, for people that haven’t read it
Alex Zhang [01:20:27]: Yeah
Vibhu [01:20:27]: What’s the one-liner for people?
Alex Zhang [01:20:29]: The simple thing is when the model is, writing its code or, like, even as, like, after it finishes writing the code, a lot of tools tend to be, like, sequential or, like, you have to wait on them, so you should just launch them in advance. Like, if you’re able to jit compile this code, you can probably figure out, like, even though it’s, like, kind of variables and stuff, like, you can figure out, like
Swyx [01:20:52]: Yeah, statically analyze.
Alex Zhang [01:20:53]: Yeah. So
Vibhu [01:20:54]: Speculation.
Alex Zhang [01:20:55]: Yeah. There is this, Someone pointed to me some actually, like, academics have, especially PL, like programming languages people, have some, like, very kind of cool ways of doing this. And so, like, at some point maybe I’ll, I’ll, I’ll, like, work on this.
Swyx [01:21:09]: I guess mostly you have to change language, because if you are in JavaScript, Python, you can’t do this.
Alex Zhang [01:21:13]: Yes. Yeah.
Swyx [01:21:14]: So, like, Haskell, yes. what’s, what’s the, what’s the normal one that’s, that’s not Haskell?
Alex Zhang [01:21:21]: Lisp.
Swyx [01:21:22]: Lisp, OCaml.
Alex Zhang [01:21:24]: OCaml, oh, yeah.
Swyx [01:21:25]: Yeah, any functional language
Alex Zhang [01:21:25]: Yeah
Swyx [01:21:25]: You can actually, like, pipeline this.
Alex Zhang [01:21:27]: Yeah.
Swyx [01:21:27]: So Effect-TS if you wanna do TypeScript.
Alex Zhang [01:21:29]: Yeah.
Swyx [01:21:30]: Okay, we can switch over to,
Vibhu [01:21:32]: I like this diagram.
Alex Zhang [01:21:33]: Yeah.
Swyx [01:21:34]: Which is like your, you guys’ whole thesis, right?
Capability Overhang: Reliable Long-Running Work
Swyx [01:21:36]: Like, that, to me this is, like, kind of like a restatement, but maybe I’re missing something of
Alex Zhang [01:21:41]: Yeah
Swyx [01:21:41]: Like, well, work on better harnesses or, like, your models actually are capable a lot more if you try harder, so this is a skill issue.
Alex Zhang [01:21:48]: Yeah. Yeah, basically. I think there’s one thing I want to see. I appreciate that there’s a big focus on, like, jagged intelligence, because it paints a big picture of, like, we can do this if we really set our minds on it. But I kind of wish. And maybe someone in academia should do this. Like, really just sit down and think about, like, if I took Astra, even the current frontier models are not good enough at, like, doing a particular job over, let’s say, the span of a month consistently and well. And I think this is, like, a stupid problem. Like, I genuinely think we can solve this. You don’t need to be a frontier lab and, like, do all this, like, fancy stuff for your IPO. Like, I think, like, these models are so smart that even if it’s, like, a silly way, I think that it genuinely is a skill issue of you can get a model to be as good as, let’s say, like, just some 18-year-old high school kid
Alex Zhang [01:22:50]: Doing some job. I think it’s, like, ridiculous that we can’t do that. And it’s. Part of the reason is, like, the format of a language model is not really amenable to that, but I think you can shape a harness around it and do it. And I think, like, this in itself is, like, ignoring the RLM stuff, ignoring all the, like, what abstractions should we use? Like, I just think someone can design a harness that can do this. Like, I, that’s. Maybe it’s
Swyx [01:23:13]: When you say do this, do what?
Alex Zhang [01:23:15]: Do long-running but simple tasks, and do them reliably.
Swyx [01:23:20]: Okay.
Vibhu [01:23:20]: What’s an example? So is this different than, like, pick your favorite company, Harvey, for example Using LLMs to do legal work, or what’s the.
Alex Zhang [01:23:30]: I guess it’s kind of like that, except if the bottleneck was not, like, certain legal knowledge or something. Like, I don’t know. Let’s say,
Vibhu [01:23:38]: I guess, like, the examples, people can take models and build pipelines or whatever and Have an agent repeatedly do whatever it has they want, right?
Alex Zhang [01:23:47]: Yeah. So for example, like, if I wanted a general system that I could kind of talk. I can talk to it like I would talk to an intern and basically just ask it to do. to explore some small thing. So maybe an example of this is, like, very silly auto research is maybe an example of this, of, like, not necessarily finding super novel solutions, but at least optimizing all of the easy parts of any problem. They often end up being over-indexed for, like, ML training and things like that. But yeah, I don’t know. maybe that’s, that’s, that’s not, like, super clear, but there is a lot of People’s general workflows where you probably could just vibe code up. Some specific harness to help you do, like automate this thing. Some examples are like automating, finding, like research papers and stuff like that. But usually people will design like a specialized agent to help them do this kind of thing. Or like they’ll vibe put a harness, like, and just run it or like their Slack bot or something. But I almost think there’s just like a standard form, like just a harness that you just plug in. Like you don’t-- It doesn’t need to be designed for finding papers or fi-- Like, you just kind of tell it, find this for me, and you like plug it into that setting. What I’m getting at is that I think there’s a lot of easy things that can be automated. And
Vibhu [01:25:13]: Is this like a hypothesis or a point around like capability overhang? Like even if we paused, there’s still a lot of impact to be had with current state of models?
Alex Zhang [01:25:22]: In some sense, yes. Like, I guess what I’m presenting is the easiest form of this. But what this is kind of saying is that, like, we have jagged intelligence on a lot of things. Like, for example, models are like disproportionately good at coding and math. This is saying like we can translate those abilities to many other things. So like for example, if you took-- Usually if you take someone who was an IMO gold or something and you kind of apply them to a lot of different problem-solving domains, they can figure it out. I don’t actually know if this applies to models. For example, like in GPU code optimization, one very interesting question is whe-- if you were to take out all of the GPU programming data, like from a model, but it was re-- it was like as good as Astra is now, just without, like with that taken out, would it be able to still optimize GPU kernels? like would it be able to learn in context roughly what it needs to learn and then like have some pipeline or like come up with some solution to solving like optimization tasks? And I think like there’s like a mismatch between, like if you took a human that was as smart or like knew as much as Astra, there is a mismatch between what that human can do and what Astra can do, maybe around a harness. and I think we can actually approximate the human a lot more.
Continual Learning and General Problem Solving
Swyx [01:26:42]: To me, it sounds very approximate to the continual learning problem. I think you’re
Alex Zhang [01:26:46]: That’s the best example. Yes.
Alex Zhang [01:26:48]: Yes.
Swyx [01:26:48]: Why didn’t you just say that then?
Alex Zhang [01:26:49]: Yeah. I guess I
Vibhu [01:26:50]: I was like, I was like thinking, could I just blurt out some words like learning?
Alex Zhang [01:26:54]: I’m careful with that. But yes, I
Vibhu [01:26:57]: You have a very eye.
Alex Zhang [01:26:58]: Maybe, Maybe, like a bit of a, I don’t know, like a
Swyx [01:27:03]: No. So I think you’re a very, yeah, I don’t know, I don’t know your undergrad actually. Are you like a math person generally, or
Alex Zhang [01:27:09]: A little, yeah. Yeah.
Vibhu [01:27:09]: Did you study math?
Alex Zhang [01:27:11]: Yeah, that’s what I wanted to do at least.
Swyx [01:27:12]: Like a category theory type of, abstraction where you think in categories and then you have to like then translate down to the specific. And, but then you like actually really care more about the category.
Swyx [01:27:24]: And like that’s the communication error because like everyone’s listening for the specific, but actually trying to also, convey the general.
Alex Zhang [01:27:31]: Yeah.
Swyx [01:27:32]: Which is hard. I don’t really know. you can, maybe use like a shorthand of like, “Okay, I’m at level two and then I’m gonna go up to level three Then come back to level two.” That we should have say some like epistemic, like shorthand for like this kind of thing.
Alex Zhang [01:27:45]: Yeah.
Swyx [01:27:45]: Because it’s hard. Like you’re, you’re compressing a lot into word after, like sequential word decoding.
Swyx [01:27:51]: Should we convert to Neuralese? is there like a, a better form of output than English or, Python or JavaScript? I don’t know.
Neuralese, Programming Languages, and Diffusion Thinking
Alex Zhang [01:28:06]: Yeah.
Swyx [01:28:06]: This is very kind of like a s**t post, but like people have speculated about like what is the native language that people want to out-- that models want to output?
Swyx [01:28:14]: Some people say binary. That’s, that’s Marc Andreessen’s thing. I don’t know. Yeah.
Swyx [01:28:19]: PTX?
Alex Zhang [01:28:20]: Let’s say a mix of English and Python. And I only say this because The capability of a model is somewhat a reflection of what we train them on. So we still want like. Yeah, I don’t really buy the binary argument. I guess I, like, I understand, but it’s like
Swyx [01:28:40]: Yeah. You wanna model the world in some way. I think the,
Alex Zhang [01:28:43]: Yeah.
Swyx [01:28:43]: One thing I’ll, I’ll bring up is always, which I always do in this kind of conversation is Sapir–Whorf Which is you, if you choose English, you will have locked into however long English has been around, which is, let’s say five hundred years, which is not that long. Like actually Like what you, the language that you speak constrains how you think.
Swyx [01:28:59]: And if you learn a different language, for example, someone, in Chinese, we don’t have tenses. I don’t know if you. I actually didn’t know that. And I speak Chinese.
Alex Zhang [01:29:09]: Oh. I did know that, but my Chinese is not great.
Swyx [01:29:12]: Okay.
Alex Zhang [01:29:13]: Yeah.
Swyx [01:29:13]: Yeah. Or like, in, let’s say in Japanese or, I know, I forget what language it is. Like in Korean, everyone you speak to, yeah, you have to like acknowledge social status.
Alex Zhang [01:29:23]: Yeah.
Swyx [01:29:24]: But it’s a different dimension than gender, right? Like, it just like, it just influences everything you do. when I take Ling 101, apparently there’s a, there’s a, there’s a language in Africa where like there’s a vegetable gender. yeah, right? Like just like you have. Or like Eskimos know no word for snow or whatever. Like, anyway, so like the language that you adopt affects your thinking. And if you Choose to output your chain of thought in English, you are biasing towards whatever English solves. I don’t know what the sort of prior of English is.
Alex Zhang [01:29:51]: That’s interesting. I did not think of it that way.
Vibhu [01:29:55]: At some level it’s interesting, right? So you’re right on language. a lot of model chain of thought also fluctuates language.
Vibhu [01:30:03]: The obvious example is Chinese models Speaking in English might still reason in Chinese. but at the same level, most models are very capable multilingually.
Vibhu [01:30:13]: And that adaptation we can see, you can add in languages. You don’t get that much from adding a whole language, but
Swyx [01:30:18]: Yeah.
Vibhu [01:30:18]: They will reason interchanged
Swyx [01:30:21]: Yeah. So we’re all autoregressive. but also like
Vibhu [01:30:23]: Right.
Swyx [01:30:23]: Let’s say German, like, subject-object, agreement, you have to put the verb at the end, which is very super annoying, like very famously. yeah, right. You don’t know what you’re doing until the end where you’re like, “Oh, that Mess of nouns and then the verb.” Well, the most classic one that most people be-- have heard of is Arrival, where, they have the heptapods where they think, that time is like flat to them. So they think in, they output entire sentences at one shot. so it’s, this is closest to, like, the difference between autoregression and diffusion.
Swyx [01:30:55]: We talk in autoregression. What if you could talk in diffusion Where things just resolve over time?
Alex Zhang [01:31:01]: I see. I see.
Swyx [01:31:02]: But like the whole idea shows up at once.
Alex Zhang [01:31:05]: I see.
Swyx [01:31:07]: So that is a drastically differenting, language, but it is a language.
Alex Zhang [01:31:10]: I see. Oh, that’s really interesting. That’s really interesting.
Swyx [01:31:12]: Which, like, machines could speak, that we, probably will never speak, but like, yeah, machines don’t care.
Alex Zhang [01:31:19]: Maybe this is a huge tangent
Swyx [01:31:20]: Yeah
Alex Zhang [01:31:20]: But are there not, like, things inherently that are reasoning chains that are inherently autoregressive?
Swyx [01:31:30]: Yeah, time.
Alex Zhang [01:31:30]: Sometimes, like. Yeah. Or like, yeah.
Swyx [01:31:32]: Yeah. Something happens first, then something else happens.
Alex Zhang [01:31:34]: Even like, yeah, anything in code, for example, like has to be causal in some-- usually at least has to be causal.
Swyx [01:31:40]: Well, no. so it’s a. Then you have to. Then you’re not exploring enough,
Alex Zhang [01:31:44]: That’s true. Yeah
Swyx [01:31:45]: Programming language theory, where, everything is like pure functional and like completely relational And, you sort of abstract away the solver that translates the relationships that is, are always true into code. So I, yeah, I feel like this is maybe a little bit too out of my depth.
What Comes After RLMs?
Alex Zhang [01:32:01]: No, it’s interesting though. Yeah.
Swyx [01:32:01]: But I love languages Whether it’s coding or human, and I do think a lot about how that affects reasoning and the boundaries of what we can do. I don’t need to go too much beyond that. I don’t know if you have any other thoughts. my closing question was gonna be, you have all these research, directions that you wanna do. You’re, you had a GPU mode phase. You had a, RLMs phase.
Swyx [01:32:25]: Presumably you have other stuff planned, which is why you’re not, doubling down on that. By the way, I notice that it is interesting how you guys do start with the GPU side, and then you migrate towards the zero gradient side, it’s, which is what Shenyou called it. Doesn’t that feel less legit than messing with GPUs?
Alex Zhang [01:32:44]: Yeah, I guess in the sense that, like. So you did bring up that, like, I like to think about things in, like, a math-oriented way.
Swyx [01:32:51]: Category, yeah.
Alex Zhang [01:32:52]: And it’s, like, very uncomfortable sometimes to be working on, like, harnesses and agents because it’s so.
Swyx [01:32:57]: Because you think all harnesses are the same.
Alex Zhang [01:32:58]: Yeah. So it’s like super fuzzy.
Swyx [01:33:00]: So like, you just, like, two new ideas in harnesses. Got it.
Alex Zhang [01:33:01]: Yeah. It’s, it’s It’s, it’s also just, like, empirically it’s hard to, like, verify a lot of findings, at least with the compute that we have available to us. But the reason why I think I’ve moved on to a lot of these problems is I think actually this is where most of, like, the innovation is yet to happen. To me, like, the GPU level is a means to exploring other ideas. Like, you want to, for example, like, get good at writing kernels or, like, even automate writing kernels for the sake of a broader goal of, like, I want to explore ideas where I’m not bottlenecked by systems challenges. In that sense, like, I guess a lot of what’s written there is all harness stuff, but I am also interested in things at the model level as well. but I’ll just leave it at that.
Swyx [01:33:48]: Okay.
Alex Zhang [01:33:48]: Yeah.
Swyx [01:33:48]: That’s a good hint. anything, any. If people wanna reach out to you, what are you looking for help on? What do you want collaborators on? any sort of calls to action?
Collaborating on Research and Choosing Big Bets
Alex Zhang [01:33:58]: Yeah. So, I guess there’s nothing I have in particular where I feel like I need to work with someone on, unless it’s, like. unless it’s with a company before, like, compute or, like, with. to talk with other people about it. But I will say I’m not. I’m never opposed to working on ideas with other people. I get reached out to a lot by often undergrads or even, like, other students.
Swyx [01:34:22]: Podcasters.
Alex Zhang [01:34:23]: Podcasters. and usually I feel like I get an email that’s something along the lines of, like, “I really like RLMs.” Like, “I wanna work together.” and I feel like I
Swyx [01:34:33]: Yeah, that’s a bad reach out, right?
Alex Zhang [01:34:35]: Yes, yeah.
Swyx [01:34:35]: The worst is like, “Can I pick your brain?”
Alex Zhang [01:34:36]: Yeah.
Swyx [01:34:37]: And like, “On what?” Like, “Read my paper, dude.” Like.
Alex Zhang [01:34:39]: They’ll like, they’ll be like, “I read your paper,” in quotes, like, “Recursive language models,” or like, “Prime Agent,” like a self-improving RLM harness or something. and it’s kinda like I like. I really like people that are opinionated, even if we disagree. I think if you have strong opinions and are able to, like, think through why you think those opinions are right or wrong, ‘cause usually it’s, it’s hard to actually tell. But, like, you have strong convictions about certain problems. Like, I’m, I’m always happy to, like, chat and even, like, potentially work on something together. I have, like, no limit to who or, like, what I would like to work on. So, yeah.
Swyx [01:35:12]: No limit?
Alex Zhang [01:35:13]: Yeah. I. in the era of agents, I think there’s a lot more work you can do, like, bandwidth-wise. So I. Yeah. I think in general, like, I am not hard to impress, but I think it just takes a little bit of effort to
Swyx [01:35:30]: Yeah
Alex Zhang [01:35:30]: Kind of. Yeah, know
Swyx [01:35:32]: Yeah
Alex Zhang [01:35:32]: Know what you want.
Swyx [01:35:32]: It’s very clear. and like, when you see a new thing come out, well executed, good, simple idea, then, like, get that
Alex Zhang [01:35:40]: Yeah
Swyx [01:35:40]: Immediately gets your attention, right?
Alex Zhang [01:35:41]: I get excited. Yeah.
Swyx [01:35:42]: It’s actually, like, not that hard to get the same attention that all the Frontier Lab guys
Alex Zhang [01:35:45]: Yeah
Swyx [01:35:45]: Because they are looking for you. you just have to put yourself out there, right?
Alex Zhang [01:35:48]: Exactly, yeah.
Swyx [01:35:49]: But yeah, it’s true. I will say, I think human attention very scarce right now, and I do struggle with, like, the number of projects I have going on.
Swyx [01:35:57]: And, I don’t know how to manage it. I don’t think agents are helping at all.
Swyx [01:36:00]: Like, I will just prompt it and. I’ll prompt a thing and then never look at it.
Swyx [01:36:03]: Right? Like, which is very common.
Alex Zhang [01:36:05]: Yeah.
Swyx [01:36:05]: Yeah, and that sucks.
Alex Zhang [01:36:06]: I guess, maybe the. one of the smaller differences in, like. actually, maybe you were doing research. I’m not sure. But for me at least, like, I’ll have maybe, like, 10 or 15 different ideas that I wanna do, but the thing is, like, most of them are bad.
Swyx [01:36:20]: Yeah.
Alex Zhang [01:36:20]: And this also maybe is true even for someone that reaches out to me. Like, maybe the idea is actually bad, but it looks interesting to me. And so, like, we can spend, like, some time looking into it, and if, like, we feel like there’s actually something there, like, then we’ll. we should take the next few weeks and just really pursue it. And, like, this is my style with. This is why I love the PhD, by the way, because there are times when I’m just thinking about problems, like, maybe on a run or just, like, playing tennis or something. Like, I’m not, I’m not working, I guess. But it’s like those are the most fun times, and then when I, like, really am convicted about something, I’ll just, like, drop everything and just do it.
Alex Zhang [01:36:52]: Like, just spend, like, all my time thinking and working on that problem. And then, once you get to the point where, like, you can just run experiments, then it’s, it’s, it’s kind of easy coasting again, so.
Science as the Next Frontier
Swyx [01:37:03]: Yeah, sorry. This is
Alex Zhang [01:37:04]: Yeah. No problem
Swyx [01:37:04]: Like the, for the fourth last question, which is like, I think a lot of people are also thinking about science as the next frontier, like physical sciences Bio, math even. how do you separate, like, I guess, let’s say your choice of projects that is applicable for industry And then maybe it’s part-- choice project is just, like, science?
Alex Zhang [01:37:27]: I actually worked on, like, AI for bio stuff before I started my PhD. The field has changed a lot since then
Swyx [01:37:33]: Yeah
Alex Zhang [01:37:34]: I should say.
Swyx [01:37:34]: ‘Cause it’s, it’s like, it used to be a theoretical, like, of course, what do you mean? Like, I have one path and then
Alex Zhang [01:37:39]: Yeah
Swyx [01:37:39]: I chose that. But now a lot of people are crossing over.
Alex Zhang [01:37:41]: Yeah.
Swyx [01:37:41]: And like, so we have started a science pod to just Cover those things
Alex Zhang [01:37:44]: Oh, wow
Swyx [01:37:45]: Because a lot of engineers are like, “Well, actually there’s, Tractable problems there.”
Alex Zhang [01:37:50]: Yeah. I will preface by saying my understanding of a lot of these topics is probably pretty limited. But I think, like, if I find out either because someone reaches out or, like, I look at a problem and I’m like, “Hey, like, some design principles that we use or that we’re thinking about right now actually make a lot of sense for this problem,” I get excited about those as well. But I think it’s harder. I don’t know. I think with. I think science, especially like empirical or, like, applied science has very long, like, what is it called? Like, feedback loops or whatever.
Swyx [01:38:23]: Yeah, it converges to a robotics question.
Alex Zhang [01:38:25]: Yeah. to me, like, also this aspect of, like, what is worth spending and betting my time on now? Because, like, maybe I spend a lot of time on this problem and then, like, in six months, like, a different solution, kind of like maybe a new model comes out and it’s like, “Oh, it’s way better for this.” And so I do have to be careful. I, like, you have to be conscious about, like, where you think things might be going.
Swyx [01:38:47]: Yeah, exactly.
Alex Zhang [01:38:47]: So, yeah.
Swyx [01:38:48]: Publish cycle.
Alex Zhang [01:38:49]: Yeah. So
Vibhu [01:38:51]: Which then I can just tell ARC-AGI-3, it’s saturated. We did it.
Alex Zhang [01:38:54]: Yeah, ARC-AGI-3 got saturated in less than a year, so it’s kind of ? Like, it’s. I don’t know. Like, if you were a lab picking that problem, like, you’re probably kind of sad now ‘cause actually
Swyx [01:39:04]: Yeah. Well, so exactly. That’s why knowledge work, gaming, all these things are saturated.
Alex Zhang [01:39:08]: Yeah.
Swyx [01:39:08]: Now they’re actually the frontier is science. So
Alex Zhang [01:39:09]: Yeah.
Vibhu [01:39:10]: Knowledge work is saturated.
Swyx [01:39:12]: Yeah. GDP val is like 80 something, 90 something.
Swyx [01:39:16]: Like, there’s, there’s 90 to 100% that was obviously gonna get
Alex Zhang [01:39:20]: It’s gonna be really hard to tell
Swyx [01:39:21]: 10 years. But, like, well, the next low-hanging fruit is gonna be
Alex Zhang [01:39:25]: It makes sense
Swyx [01:39:25]: The other stuff.
Alex Zhang [01:39:26]: Yeah. Maybe I’ll think about that more. I actually, I haven’t given too much thought to
Swyx [01:39:30]: I’m just trying to guess your next direction, actually.
Alex Zhang [01:39:32]: No. I will say because I think especially at MIT, like, it’s, it’s, it-- Yeah, there’s a lot of really talented scientists there, like, in the natural sciences. And I think it’s, it’s a little bit, like, sacrilegious almost to be like, “I’m gonna figure out, like, your problem.” Like
Swyx [01:39:45]: No, that’s how, that’s how it’s done.
Vibhu [01:39:46]: It’s great.
Alex Zhang [01:39:46]: Oh, no, I know. Yeah.
Swyx [01:39:47]: So I interviewed Yitai who did the IMO thing. He’s just like
Alex Zhang [01:39:50]: Oh, yes. Oh, yeah.
Swyx [01:39:50]: “I’ve never, I’ve never been to IMO. I don’t even know what it is.” It’s a skill model, dude.
Vibhu [01:39:55]: Yeah.
Swyx [01:39:56]: Which is, like, very disrespectful, but like, whatever. But that’s why, yeah.
Vibhu [01:40:00]: Yeah. At some point, like, you have to respect, like, okay, the progress is being made.
Vibhu [01:40:05]: Like, number is getting output, right?
Alex Zhang [01:40:07]: True. Yeah, true.
Swyx [01:40:08]: Yeah. that is a very big lesson. It’s very interesting, like, ‘cause the mathematicians are responding this way to NARI systems right now.
Alex Zhang [01:40:13]: Right. Yeah.
Swyx [01:40:14]: Like, Terry Turnstow is like
Vibhu [01:40:15]: Turnstow
Swyx [01:40:16]: “No, like, let’s not, let’s not use AI.”
Alex Zhang [01:40:17]: Yeah.
Swyx [01:40:17]: I’m like, “Mm, I don’t know.”
Alex Zhang [01:40:19]: Well, yeah. I think that whole thing is kind of weird ‘cause I feel like, I feel like they would have had a stronger case if a lot of them didn’t work with OpenAI before, like, all this happened.
Swyx [01:40:30]: No, that’s ad hominem, and they’re really trying to stay away from that. Then so what, right?
Alex Zhang [01:40:34]: So what? Yeah.
Swyx [01:40:35]: Like, I don’t know. Like, so what? They got. they’ve, they’ve collaborated. I collaborate with people that I don’t
Alex Zhang [01:40:39]: I guess it’s true
Swyx [01:40:39]: I don’t agree with or
Closing: Research, Academia, and What Comes Next
Alex Zhang [01:40:40]: That’s true
Swyx [01:40:40]: Whatever.
Alex Zhang [01:40:41]: Yeah. That’s true.
Swyx [01:40:41]: Or, like, I did a thing and then now I regret that. I changed my mind. whatever.
Alex Zhang [01:40:45]: Yeah.
Swyx [01:40:45]: So I’ll defend their right to say that. But like, yeah, a lot of people are reasonably disagreeing with them.
Alex Zhang [01:40:50]: Yeah.
Swyx [01:40:51]: Okay, cool. thanks for your joining us. Congrats on, your success so far. I can’t believe you’re still not done with your PhD.
Alex Zhang [01:40:58]: Well, it’s year two?
Vibhu [01:41:00]: Yeah. Can’t believe we did this podcast without going through the RL paper.
Swyx [01:41:05]: He had a, he had a definition.
Alex Zhang [01:41:07]: I think the paper is more about, like, empirical results. Like, the actual idea is quite simple.
Vibhu [01:41:11]: Yeah. And you’ve talked about it many times.
Alex Zhang [01:41:13]: Yeah, at this point. I think there’s, there’s more interesting things to look over now, so.
Vibhu [01:41:17]: Cool.
Alex Zhang [01:41:17]: Yeah.
Vibhu [01:41:18]: Well, we’re excited to see what you do next.
Alex Zhang [01:41:19]: Thank you so much.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week
30/09/2026 | 39 minThree months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks:
We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch, there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein, cofounder of Sky and now leading all the amazing CUA progress that casuals might miss:
Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software.
OpenAI clones Jev
In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API. Given that we were the first Jev podcast, we particularly focus on the unusually fast sprint on the Decisions API:
And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns.
We discuss:
* Why OpenAI thinks Computer Use has changed dramatically in just the last few months
* Dots and what changes when every agent gets its own Linux computer
* Why Computer Use can now complete some tasks faster than the average human
* The path from human-level to “literally superhuman” computer use
* Why modern agents are much better at debugging and recovering from failure
* How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together
* App Shots and why they give models much richer context than ordinary screenshots
* Why Computer Use can close the loop between writing software and testing it
* Trust, permissions, and safety when agents can make payments and operate websites
* Async function calling and why models no longer need to stop reasoning while tools run
* Mid-turn steering, WebSockets, and the architecture behind more responsive agents
* UltraFast inference and how OpenAI is pushing frontier models toward much lower latency
* The rapid internal story behind the Decisions API
* Why Decisions API is more than structured outputs at low latency
* GPT Live, fast tool calling, and real-time computer control
* How OpenAI is already using Decisions API for support classification and internal workflows
* Longer prompt caching, cache pre-warming, and cache-aware applications
* Server-side compaction vs manual compaction for long-running agent threads
* What should live inside an Agents API versus a developer’s own harness
* OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs
Ari Weinstein
* Product & Engineering, Computer Use at OpenAI
* X: https://x.com/AriX
* LinkedIn: https://www.linkedin.com/in/weinsteinari/
Nikunj Handa
* Product, API at OpenAI
* X: https://x.com/nikunjhanda
* LinkedIn: https://www.linkedin.com/in/nikunjhanda/
Timestamps
00:00:00 OpenAI DevDay: Dots, GPT-6.1, Agents API, and Decisions API
00:02:52 Dots and Personal Cloud Computers
00:04:59 Why Computer Use Is “180 Degrees Different”
00:06:04 From Sky to Self-Debugging Computer Use Agents
00:09:24 How Computer Use Sees and Operates Software
00:12:09 From Faster Than Humans to Superhuman Computer Use
00:16:03 Agents API: Trust, Permissions, and Safety
00:17:31 Computer Use for Coding, Testing, and QA
00:19:14 GPT-6 APIs, Async Tool Calling, and UltraFast Inference
00:23:21 The Rapid Story Behind Decisions API
00:25:32 What Decisions API Is and How It Works
00:30:24 What OpenAI Is Building With the New APIs
00:32:23 Prompt Caching, Pre-Warming, and API Performance
00:35:20 Context Compaction for Long-Running Agents
00:37:13 Memory, Higher-Level APIs, and the AI Cloud
Transcript
Introduction: OpenAI DevDay and the New Agent Stack
Vibhu [00:00:00]: Okay. We’re very excited to be here. Today is OpenAI DevDay. Special podcast
Swyx [00:00:08]: We’re the first podcast after your livestream.
Vibhu [00:00:10]: First podcast. We have Ari here, who leads the product and engineering team for Computer Use agents. Before we kick in and dive deep on Computer Use, you wanna give a quick recap? What was announced? What’s the quick slew of announcements you guys had today?
Ari Weinstein [00:00:24]: Yeah. yeah, it was a super exciting day. we just got out of the keynote. It was really sick. there were a bunch of Computer Use announcements that I think are worth thinking about. We have, Dots, which is the new, sort of personal assistant product, and, that has some really exciting Computer Use features. There’s GPT-6.1 Sol, which is this amazing new model, that I think is particularly great for Computer Use ‘cause of, sort of the cost and speed, advantages. I think, I think we shared that it’s, a fifth of the cost of Astra and a seventh of the cost if you’re looking at Computer Use specifically, which is really amazing. sorry, there were so many things. I’m trying to sort through it.
Swyx [00:01:02]: And the API.
Ari Weinstein [00:01:03]: Agents API, which now has Computer Use in it, which is really cool, ‘cause now developers can build on the same Computer Use, that is part of Codex, and ChatGPT. and then there were some demos of our existing Computer Use features, like app shots, where you can take the context of something you’re doing on your computer and bring it into Codex and ChatGPT really fast. And then, like, native Computer Use on your Mac, where Roman had it taking screenshots of his app, automatically, and he could do other things on his computer while Computer Use was using his applications. so yeah, really exciting keynote.
Swyx [00:01:35]: And not to mention the Decisions API.
Ari Weinstein [00:01:37]: Decisions API.
Swyx [00:01:38]: Off the bat, are they all the same model? Like, this is. Or the same dataset distilled to different models?
Swyx [00:01:44]: Like, basically, like, is Computer Use using Decisions API, or are they, like, kinda separate?
Ari Weinstein [00:01:49]: So what’s really cool about the Decisions API is it, you know, it has all these new capabilities. It does inference in parallel. it doesn’t have reasoning. It’s a smaller model, than the ones we use for Computer Use. and so those capabilities make it really fast.
Dots and Delegating Work to a Cloud Computer
Swyx [00:02:07]: Yeah.
Ari Weinstein [00:02:07]: They also make it a little bit less good at doing, like, long horizon, sort of sophisticated tasks. And so I think I would say it’s still an open area of research for how we, like, bring those approaches together. But, yeah, I’m really excited to see what people build with the Decisions API.
Vibhu [00:02:24]: One of the interesting things is Dots now have attached personal computers.
Ari Weinstein [00:02:28]: Yeah.
Vibhu [00:02:28]: So it seems like they’re very much more persistent. You’ve been using them for a while. How should people push the bounds? Like, what should people aim for? What should they try? Personally, right now I use it for a lot of customer service. Like
Ari Weinstein [00:02:41]: Cool
Vibhu [00:02:41]: “Oh, this was wrong. I don’t wanna sign in. I don’t wanna authenticate.” Find whatever and just get it fixed.
Ari Weinstein [00:02:45]: Yeah.
Vibhu [00:02:46]: How should we push further? What should people try?
Ari Weinstein [00:02:50]: Dots Are a really cool product because each Dot has access to its own Linux virtual computer in the cloud, which is different from our other products. you know, traditionally, we’ve have access to a browser in the cloud, or it has access to your own computer, but now you get your own entire Linux computer in the cloud. And so it can run full desktop applications, and it can also use a web browser. And so, yeah, you know, I think the powerful thing about Computer Use and the reason why I think it’s so, exciting is because it makes it so that the agent can do anything you as a, as a person can do, because all the software in the world was designed for humans, and now agents can use that same software, and you can delegate to the agent. So, yeah, like, anything that you would do on a computer, you can ask a Dot to do. Yeah, I think what particularly is useful is gonna really depend on who the end user is and what- what’s valuable in their life. but yeah, I would just start by thinking about, like, one of the things that you spend time on and how could you delegate those to an agent.
Swyx [00:03:47]: Yeah, a lot of flight booking and shopping and honestly even, like, playing a game or whatever, right?
Ari Weinstein [00:03:52]: Totally.
Swyx [00:03:52]: Yeah.
Ari Weinstein [00:03:53]: Yeah, I don’t know. For me, something I did recently, I’ve been working on. I’ve, subscribed to a meal prep service ‘cause I was trying to, like, eat healthy, you know? And I really like this meal prep service I found because it lets me customize the meals I order to, like, a high degree of granularity. So I can say, like, “I want this many grams of chicken and this many grams of rice.” but it was so complicated. It took me two hours to do an order, and I found that I could ask Computer Use to do it for me, and it did it in 15 minutes. so I actually saved two hours. it both did it eight times faster than I could, and it saved me two hours on GPT-6.1 Sol.
Swyx [00:04:32]: Yeah.
Ari Weinstein [00:04:32]: So those are the kinds of tasks that I feel like, are really powerful.
Swyx [00:04:36]: As a creator, I can tell you automatically, immediately, my number one use case is automating YouTube.
Ari Weinstein [00:04:40]: Nice.
Swyx [00:04:40]: Because, YouTube doesn’t expose a lot of things via API.
Ari Weinstein [00:04:43]: Yeah.
Swyx [00:04:43]: And you have to just put it in a VM and just, like, run it, for, like, let’s say, let’s say their AB testing feature or making community posts. None of this is available by API ‘cause they hate developers.
Swyx [00:04:53]: Anyway, so,
Ari Weinstein [00:04:55]: I’ve heard that from our developer experience team too. They use it with YouTube a lot. Yeah. It’s really awesome.
Swyx [00:04:59]: So I wanna draw for, you know. let’s say, I wanna get a little bit spicy. One of our, the leading AI podcasts, our friends, is famous for saying that Computer Use hasn’t advanced in the last two years.
How Computer Use Has Changed in the Last Year
Ari Weinstein [00:05:12]: Yeah.
Swyx [00:05:13]: Which is a very interesting statement, and I think you’re one of the best people in the world to talk about this, like, how have things have progressed, right?
Ari Weinstein [00:05:20]: Yeah. You know, they said that a few months ago, I think, and I hope they have a different perspective now because Computer Use is, like, 180 degrees different than it was.
Swyx [00:05:26]: He’s a, he’s a tough guy to impress.
Ari Weinstein [00:05:27]: Yeah, okay. well, we’re working on it.
Swyx [00:05:30]: But, you know, you worked on. You’ve, like, basically spent your whole career working on, like, some kind of computer automation, right?
Ari Weinstein [00:05:34]: Yeah.
Swyx [00:05:34]: Like shortcuts
Ari Weinstein [00:05:35]: Yeah
Swyx [00:05:35]: At Apple, and then Sky, and then, and then joining OpenAI. Can you draw, like, what your through line is for, like, what is driving you and what- you, what wasn’t possible back then maybe
Ari Weinstein [00:05:47]: Yeah.
Swyx [00:05:48]: And, like, what your sort of milestones were.
Vibhu [00:05:49]: I guess to add on to that as a follow-up question, what’s the major change from using Codex Computer Use from, like, last week
Ari Weinstein [00:05:57]: Yeah
Vibhu [00:05:57]: Through to today? Is it model? Is it dots? Is it harness? So all the history plus what really just changed in today’s announcements?
Ari Weinstein [00:06:04]: Yeah. On the through line, I guess I’ve always been excited about automation and helping people automate tasks because then you can, like, save time in your life and focus on things that are more important to you than, like, operating a computer very intricately. And so, yeah, that was why we worked on some of those products. I was at Apple before. we made a company called Sky. we ended up joining OpenAI, which is really exciting. and I think something that was
Swyx [00:06:27]: And almost like you have to hack around Apple until Apple was like, “Fine, like, we’ll just hire you and you can just work on the inside,” right? Like.
Ari Weinstein [00:06:35]: It was, it was a cool place to get to work. what was really interesting looking back at Sky is we were, we were working on Computer Use there as well, and the models were so much less capable. And now the models, just in the last one year, have become extraordinarily capable at Computer Use. I think the biggest delta that I see is before they could, like, reliably start tasks, but then they would run into problems, and now they’re really good at debugging. They’re really good at trying again, introspecting what is and isn’t working. and I think we’ve also brought the Computer Use the Computer Use field itself has moved forward. I think we’re using more techniques. now Computer Use, often writes code. So if you actually look at it in Codex and you expand the tool calls manually, you can see that it’s not just doing one action at a time. It’s actually writing JavaScript code that it executes, that the computer executes to perform sometimes many actions at once, which is a great, you know, speed up and great capability. We use more accessibility, sort of multimodal interfaces. So, the model may use screenshots, it may use accessibility, it may use Playwright. it can use a lot of different mechanisms, based on the task at hand. and then, yeah, the model acceleration has been, has been just amazing. So, yeah, what’s different today? I think we’re making computers better all the time, so I think just, like, one day’s difference, is probably a little bit less consequential than, like, even the past month or the past two months. but, yeah, I think the Computer Use in Dot is really exciting as well as, the new model that we came out with.
Measuring Computer Use and Improving the Harness
Vibhu [00:08:03]: On the keynote, Tejal was mentioning 7x improvements in Computer Use speed, a lot better on a few benchmarks. How do you guys think about measuring it? Computer Use is one of those things where, as you say, you know, it’s improvements over time.
Ari Weinstein [00:08:20]: Yeah.
Vibhu [00:08:20]: Is it harness? Is it model? Is it post-training?
Ari Weinstein [00:08:22]: Right.
Vibhu [00:08:22]: How do you guys look at it internally about measuring how good it is, and what were the changes with the new model?
Ari Weinstein [00:08:29]: We actually have a bunch of different ways of measuring it, some of which are on different permutations and configurations of the harness. It’s a bit of a complicated story because, you know, our production products have, you know, some more safety checks, and, you know, those are configured differently based on the needs of the, of the task at hand. So there’s a lot of ways to measure it, but I think regardless of how we measure it, we find pretty consistent gains. and those gains are, sometimes in the harness and sometimes in the model. and yeah, I was really excited by this result that GPT-6.1 is even more cost-effective for Computer Use than its baseline cost improvement as compared to Astra. It’s, like, really cool to see.
Swyx [00:09:10]: Yeah. I mean, one of the visuals I really liked from the livestream was that, you’re sort of improving the Pareto frontier of, your, curve, and there was a lot of talking about how you’re improving it together with the harness.
Ari Weinstein [00:09:24]: Yeah.
Swyx [00:09:24]: Can you give some examples of aha moments that you had, whether it’s on, like, model driving the harness driving the model, whatever?
Ari Weinstein [00:09:32]: I don’t mean to repeat myself, but I think, like, introducing more modalities has been really powerful.
Swyx [00:09:36]: Okay.
Ari Weinstein [00:09:36]: One more specific example of that is, in the past, I think we saw a lot of Computer Use, products had to spend a lot of time, like, scrolling, you know? So it would, like, take a screenshot. It would try to do something. It would be like, “Oh, I gotta, like, scroll down to the next page of results,” and then it would take a screenshot, and then it would try to do something. It would scroll down again. And so I think, with accessibility and other. and, direct access to the DOM and other things like that, now the language model can actually see, like, an entire page or an entire application. It can write code that can do multiple steps at once. And so I think those have been probably the biggest single aha moments. There’s, like, a lot of tiny ones that are less exciting in comparison, but actually we do find also that a lot of speed improvements are driven by, like, a lot of little paper cuts that we gotta go in and introspect.
App Shots, Accessibility, and Better Computer Context
Swyx [00:10:21]: Yeah. A lot of really hard engineering.
Ari Weinstein [00:10:23]: Yeah.
Swyx [00:10:23]: I mean, app shots in general, right? Like, I think people don’t quite get the difference if. because there’s, like, a nice visual in Codex when it
Ari Weinstein [00:10:30]: Yeah
Swyx [00:10:30]: When you take an app shot, but they don’t maybe they get the difference that, you are able to actually drive each button and you have the, you have each text, in a very optimal representation.
Ari Weinstein [00:10:40]: Yeah. Exactly. Yeah. It’s kind of fun actually. If you wanna be, like, really nerdy about it, you can go into Codex, take an app shot by hitting the two command keys. So you grab the content from whatever app you’re working with, bring it into the, Codex or ChatGPT chat. And then the. if you click on the attachment and you click on this, like, little tiny button in the top right, you can see the raw text and you see the raw accessibility representation. And yeah, we’ve put a lot of work into, putting
Swyx [00:11:04]: Just dumping everything out. Yeah.
Ari Weinstein [00:11:05]: Dumping it out, but also making it token-efficient, doing it efficiently. There’s, like, a bit of an art to it. And, you know, it turns out that the same technology that was invented for humans, you know, who maybe have accessibility needs, who wanna use a screen reader technology, that technology is really helpful for them to be able to use computers. It’s also really helpful for LLMs to be able to use computers. So that’s been, like, really fun to get to work on.
Vibhu [00:11:27]: For context, I feel like a lot of people don’t understand app shots. They don’t even know it’s a feature.
Ari Weinstein [00:11:30]: Yeah.
Vibhu [00:11:31]: It’s when you double hit command, it pulls in what looks like a screenshot
Ari Weinstein [00:11:34]: Right
Vibhu [00:11:34]: And you’re like, “Oh, why have I opened up just a screenshot and thrown it in?” No, it’s actually pulling all the metadata, all the code, everything.
Ari Weinstein [00:11:40]: Yeah, exactly. Yeah. So it’s like, you know, if you take a screenshot of a webpage that has a link- The screenshot doesn’t include where the link goes. It doesn’t include, you know, maybe you take a screenshot of your calendar, the ca- event ti- titles are truncated, you know? But when you take an app shot, it gives, like, the language model, like, full context about everything and, that lets it, just sort of, like, do much more.
Swyx [00:12:02]: Yeah. For those who wanna see more, Jason Liu, I invited him to do a full workshop on this, at AI Engineer.
Ari Weinstein [00:12:07]: Amazing.
Swyx [00:12:08]: Did a great job.
Vibhu [00:12:09]: I have a broader vision question
Toward Superhuman Computer Use
Ari Weinstein [00:12:11]: Yeah
Vibhu [00:12:11]: On Computer Use agents. So your example of take a screenshot, scroll page, take a screenshot is where we were.
Ari Weinstein [00:12:17]: Right.
Vibhu [00:12:17]: Today, they can automate a lot. what are the bottlenecks? Is it models? Is it harnesses? What. Where do you see it going in, like, two years? Do you see it just running for hours? How do we get there? Any predictions on where Computer Use goes?
Ari Weinstein [00:12:32]: Yeah. I mean, I think what’s really crazy that I think, You know, the team’s accomplished over the past couple of months is that now Computer Use is, like, faster at accomplishing tasks than, like, the average human probably in most cases. and I think that the next frontier is to have Computer Use be, like, literally superhuman in its performance where it actually is as fast or faster at using software than, like, expert Computer Users like us. and I think that’ll be really consequential and exciting when that happens because I think we’ll be able to all of a sudden build products, that, provide just much more real-time experiences. And I think it’ll also. lowering the barrier to entry of, or the activation energy, I suppose, of using Computer Use I think will make us start to default to doing certain things in agents that we’ve become accustomed to doing manually. And I think that’s exciting also ‘cause it’ll save us a ton of time. and I think there’s a, you know, there are a lot of different little paper cuts and bottlenecks that are sort of standing in the way of that. I think that there’s, yeah, there’s things on the model side, there’s things on the inference side, there’s things on the harness side, there’s things in the, in the representation. You know, we find that as Computer Use gets faster, we’re increasingly bottlenecked by just, like, the speed of doing an operation. Like, for example, you know, a non-trivial amount of time in our benchmarks of Computer Use tasks is actually, like, let’s say you’re automating a task on doordash.com. Like, a lot of the time is actually waiting for doordash.com itself to load, you know?
Swyx [00:14:04]: Yeah, then you just write a wait and then you execute the wait.
Ari Weinstein [00:14:07]: Yeah, totally. And you wanna get. Yeah, actually, it’s actually really important that you get that de- like, you want as little delay as possible between when it finally finishes loading and when you go and
Swyx [00:14:16]: Yeah
Ari Weinstein [00:14:16]: Trigger the LLM to do the next action, which is actually- itself a statistical science.
Swyx [00:14:20]: Like an event-driven way maybe to do that.
Ari Weinstein [00:14:22]: When possible, you want it to be event-driven.
Swyx [00:14:24]: JavaScript has some load events.
Ari Weinstein [00:14:25]: And JavaScript has load events for. or the web browser has load events for web navigation, but there’s other types of events that actually really can’t be event-driven. So there’s a lot of complexity
Vibhu [00:14:34]: The one that comes to mind is, like, chatting with customer service.
Ari Weinstein [00:14:37]: Yeah.
Vibhu [00:14:37]: Replies could take 30 seconds, could take three minutes.
Ari Weinstein [00:14:39]: Oh, right.
Swyx [00:14:41]: I have dealt with so many bots with Codex. it’s great, but I also wonder if the other side knows that they’re talking to a bot ‘cause I’m, like, answering in complete sentences. Like, I’m capitalized correctly.
Ari Weinstein [00:14:50]: That’s hilarious.
Swyx [00:14:51]: Like, I’m giving full num- full reference numbers and everything. Like, it’s too. it’s clearly too good. I don’t care. Like Like, I’m just, like, trying to get my support case.
Vibhu [00:14:58]: I’ve prompted it to, like, you know, “Don’t pretend you’re a bot. Be very annoyed human.”
Vibhu [00:15:02]: Short one-liners, like
Swyx [00:15:04]: Yeah
Vibhu [00:15:04]: Push it, do all this. I also tell it, “While you’re waiting for responses, like, use subagents to research better ways to figure out what we need.”
Ari Weinstein [00:15:12]: Nice.
Vibhu [00:15:12]: It’s just, like, human little intervention.
Ari Weinstein [00:15:14]: That’s awesome. I also feel like half the time it’s a bot on the other end, so now you
Vibhu [00:15:17]: Yeah
Ari Weinstein [00:15:17]: Got the bots talking to each other.
Swyx [00:15:18]: Yeah. I will also say, you know, like, you know, one milestone of Computer Use that we are, we’re at now is, you know, three, four years ago, we were scared of hooking up LLMs to the, to the web and to
Ari Weinstein [00:15:31]: Yeah
Swyx [00:15:31]: To our, to our devices. And now I’m having it configure DNS for me.
Ari Weinstein [00:15:35]: Wow.
Swyx [00:15:36]: I’m having it pay my bills, and, like, really, like, tens of thousands of dollars of, like, stuff I’m just sending it over and yoloing with Computer Use and, like, you know, what’s the, what’s the worst thing that can happen?
Swyx [00:15:48]: So that- that’s all, that’s all really good.
Building Safely With Computer Use in the Agents API
Ari Weinstein [00:15:50]: Yeah.
Swyx [00:15:50]: I think now that you’ve. you know, obviously, you also have to dogfood your own products and all these things. Now that you’ve sort of released this in API, what are some pitfalls or tips that you wanna tell developers, because they’re about to, I guess, encounter all this, firsthand?
Ari Weinstein [00:16:03]: First of all, I’m just really excited that we brought Computer Use into the Agents API. I think this is, really great because obviously a lot of developers are building applications that wanna be able to work with third-party websites and services. And so Computer Use has this universality to it. It can work with anything. So now all of a sudden, developers can build using the same Computer Use implementation that we’re building on. I think there’s great work to be done if you wanna build your own Computer Use harness, but it’s hard. And also, we train our models on our Computer Use harness, so there is, an advantage to using the one that’s in distribution for the model. There actually might be a speed and cost and accuracy advantage. So I think it’s really great for people to get to build on top of that. And, yeah, you know, I think kind of to the point that you were making, like, I think we’re all sort of still in the process and maybe, like, some of us are ahead of many people in the world of, like, getting comfortable with this technology and trusting it. And so I think it’s incumbent on us to, sort of build that trust over time by making sure we’re building things that are reliable, by building, the right kinds of safety checks, by asking for the user’s consent before doing something consequential like making a payment, by, asking, you know, maybe depending on the application, making sure you’re only letting it access the websites or applications that it actually needs for the task. So that’s, I think, something important to think about. but yeah, I’d really encourage people to try the new Agents API, build all kinds of cool stuff on it. We’d love to hear your feed- feedback if, you know, depending on how it goes.
Vibhu [00:17:31]: Have you seen any changes in the way it affects dev workflows? So one of the things with dots is, you know, you’re seeing it in Slack.
Ari Weinstein [00:17:38]: Yeah.
Vibhu [00:17:38]: You’re seeing people use voice and build. the example Roman showed of change this app and send me screenshots along the way and all this.
Computer Use for Testing and Closing the Software Loop
Ari Weinstein [00:17:46]: Yeah.
Vibhu [00:17:46]: Is anything that you’re seeing there in adoption about how people are using Computer Use for coding workflows? Any tips people should take from that?
Ari Weinstein [00:17:55]: One of my favorite use cases for Computer Use actually, and one that we see a lot in the wild, is Computer Use letting the agent- actually test the software that the agent has built, which is far more consequential than it sounds. Because traditionally, you know, you’d build something in Codex and then the-- and the Codex builds it for you, and then you have to test it, and you are now like QA for the agent, right? So with Computer Use, you can complete the develop-- the software development life cycle, where, the agent can build software, it can test it. So I have a lot of fun, you know, building stuff, having the agent test it. By the time it comes to me, it’s already working. I have, extra fun because sometimes I’m, like, developing Computer Use itself, and so now I have a Computer Use agent that’s using my Computer Use agent that’s using something else. so yeah, I really, I really think this is a super powerful class of use case.
Swyx [00:18:44]: I have a visual play test skill that I’ve developed that, really catches a lot of design issues,
Ari Weinstein [00:18:49]: Nice
Swyx [00:18:50]: That, you know, normally when you just look at code, you wouldn’t really pick it up. it’s also really good for cloning apps, though. If you’re using a shitty SaaS and you wanna kill the SaaS You just clone it screen by screen by screen. and Obviously, Computer Use can completely drive everything, take screenshots, note it down, and then clone everything with Codex.
Ari Weinstein [00:19:06]: That’s really cool.
Swyx [00:19:06]: But yeah, thanks for all your progress. I think, that is
Ari Weinstein [00:19:08]: Absolutely
Swyx [00:19:09]: Our time.
Nikunj Handa: What’s New in the OpenAI API
Ari Weinstein [00:19:10]: Yeah.
Swyx [00:19:10]: This is not the last that we’re gonna talk.
Ari Weinstein [00:19:12]: Yeah, cool. This has been really fun. Thank you guys for having me.
Swyx [00:19:14]: All right.
Vibhu [00:19:14]: All right. Okay, we’re a strict cutoff. We’re just gonna dive right in.
Nikunj Handa [00:19:17]: Let’s do it, yeah.
Vibhu [00:19:19]: Okay, so, Nikunj, we’re very excited to have you. You shipped a lot on the API side, like we just
Nikunj Handa [00:19:25]: Yeah
Vibhu [00:19:25]: Talked about with Ari. You can now build with Computer Use agents. Anything you wanna highlight, the API side of changes, and introduce yourself a little and what you do?
Nikunj Handa [00:19:34]: Yeah, for sure. My name is Nikunj. I lead product for the API team. Been here for roughly three years. been working on launching models. I feel like that’s just been, like, a thing, constant thing throughout my time, here at OpenAI. And, with every new model, we try to, like, basically work super closely with the post-training team, the research team, to figure out what’s new in it. and then we, like, expose those capabilities in the API. so that’s, like, the basic way of putting it. and if you just look at, everything that’s new with GPT-6, the cool new capabilities that we launched were, firstly, async function calling. so what you see with, like a lot of the things that you’re seeing in, like, Codex and Dots and everything is that tool calls take so long that you don’t have to, like, pause the model’s execution while, the tool is running. So you could just, like, kick off a tool call, keep running, keep reasoning, and then check back in. so we launched async tool calling. We launched, like, mid-turn steering, so now you can, like, inject messages while the model is reasoning, in the middle. so as your tool call finishes, you can put in that instructions.
Async Tool Calls, Mid-Turn Steering, and WebSockets
Swyx [00:20:43]: And that’s also partially a model alignment capability, right?
Nikunj Handa [00:20:46]: Yeah.
Swyx [00:20:46]: Like, they have to train in the ability to train.
Nikunj Handa [00:20:48]: Exactly, yeah. And
Vibhu [00:20:49]: I feel like we’ve had it in the app. You could always, as it’s reasoning, you could steer.
Nikunj Handa [00:20:54]: Yes.
Vibhu [00:20:54]: It wasn’t the best. It’s gotten much better.
Nikunj Handa [00:20:57]: Yeah.
Vibhu [00:20:57]: Excited to see how it does this in version
Nikunj Handa [00:20:58]: Yeah, and I like our main
Vibhu [00:20:59]: And now
Nikunj Handa [00:21:00]: Goal in, our main goal in the API is to, like, put things in the API once it’s trained into the harness. And so we kinda wait for that moment until it’s good enough. And a lot of that is, like, actually being powered by WebSockets, which we launched, a few, I wanna say months ago. And so WebSockets just opens this, like, whole bidirectional, like, communication thing with the model. This is not, the GPT Life thing. I’m just talking about GPT-6. and you can do all these, like, async tool calling, async reasoning, injecting messages. It’s a really fun API to work on. I think, like, really enjoying.
Swyx [00:21:33]: Yeah. This is why we are the engineering podcast, because we get to talk about WebSockets.
UltraFast and the Inference Stack
Nikunj Handa [00:21:36]: Yeah.
Swyx [00:21:37]: This also pairs very well with UltraFast, right?
Nikunj Handa [00:21:39]: Oh, yeah.
Swyx [00:21:39]: Like, that is now, like, I think for the first time ever available in the API.
Nikunj Handa [00:21:43]: Yes.
Swyx [00:21:43]: Which is, which is basically the theoretical fastest speed you can ever get, Frontier of Intelligence.
Nikunj Handa [00:21:49]: Yeah. It’s been so exciting to work on that project. I think, before I go into the API, the most fun part of, UltraFast has been just watching the inference team cook with Astra. Like, they’re just, like, constantly having these, like, Codex agents running, trying to, like, squeeze out more performance. And, I would say, like, at least for a couple of months, a lot of it was focused on efficiency and driving the cost down, which is how we, like, were able to cut the Luna price by, like, 80%. It was, like, a lot of that was driven by, like, all the inference improvements they landed. And then now they’ve, like, shifted gears towards, like, how can we make this run as fast as possible? And so UltraFast has just been, like, amazing to see on a mo- on a model like Astra. Like, to go that fast has been really cool. And yeah, WebSockets is like. actually it was like the first time we launched WebSockets, it was for GPT, 5.3 Codex Spark, which was. Can’t believe we named a model that, but, you know, that’s what we launched it for. And obviously, it helps so much because, like, you gotta have the tool calls. you had, like, really reduced the overhead, of going back and forth with tools. And so, WebSockets is awesome for that.
Swyx [00:22:57]: Yeah. it’s always cute to see, like, I have my reset usage limit, and then I have my Spark usage limit that I never use.
Nikunj Handa [00:23:03]: Yeah.
Swyx [00:23:04]: Like, it’s there if I want it.
Nikunj Handa [00:23:05]: I think it’s gone finally.
Swyx [00:23:06]: It’s gone. It’s gone, yeah.
Nikunj Handa [00:23:07]: I know it’s gone, so.
Swyx [00:23:08]: Yeah. you’re slowly killing off all the, you know, the
Nikunj Handa [00:23:11]: The old ones, yeah.
Swyx [00:23:11]: Oldies.
Vibhu [00:23:11]: This is a great week. I mean, it was the first time we had Frontier Intelligence at extreme speeds.
Nikunj Handa [00:23:17]: Yeah.
Vibhu [00:23:18]: People really liked it.
Nikunj Handa [00:23:19]: Yeah.
Vibhu [00:23:19]: So
Swyx [00:23:20]: Yeah
Vibhu [00:23:20]: First time it comes back.
Swyx [00:23:21]: Yeah. for, 5.3 Spark is explicitly attributed to Cerebras. You guys are not confirming or denying that, UltraFast is related to Ce- Cerebras, but people are. I’ll just say that people do care and, are wondering about it. And you have your own silicon as well. elephant in the room, decision models.
Decisions API: OpenAI’s Fast Decision Model
Nikunj Handa [00:23:38]: Oh, yeah.
Swyx [00:23:38]: Decisions API. We were the first podcast to do a big Jev, deep dive with, Diogo, and I also, you know, featured him at AI Engineer. How quickly did you see Jev and go like
Nikunj Handa [00:23:49]: Oh my gosh. Yeah.
Nikunj Handa [00:23:50]: Yeah. Firstly, like, huge props to Diogo and, like, the Jev team for, like, really inspiring the
Swyx [00:23:55]: Yes
Nikunj Handa [00:23:55]: Like, whole segment in the market. Like, obviously Jev comes out, everyone’s, like, losing their minds over it. Our users are, like, hitting us up. But also, like, our internal teams are like, “We need, like, a much faster classification system.” We can. I don’t wanna, like, get ahead of some of the dots features that are gonna come
Swyx [00:24:16]: Whoo
Nikunj Handa [00:24:16]: But you’re gonna see, like, some cool, like, really snappy, fast things built on top of the decisions API. but, you know, like, yeah. Props to Jev for, like, inspiring this whole thing. obviously a bunch of people at OpenAI get nerd sniped by that, and they’re like, “How can we, like, make this work? We’re not gonna, like-”
Swyx [00:24:33]: Okay.
Nikunj Handa [00:24:33]: “. train a new model.” But
Swyx [00:24:34]: Like, four weeks ago, this was not on the dev radar, right?
Nikunj Handa [00:24:37]: No, not at all. No.
Swyx [00:24:37]: Okay.
Nikunj Handa [00:24:37]: This is like
Swyx [00:24:38]: Wow
Nikunj Handa [00:24:38]: Jev-inspired and, like
Swyx [00:24:40]: I think you are officially the first one to your lab to, like, clone and, adopt this.
Nikunj Handa [00:24:44]: Yeah. Yeah. I feel like, OpenAI has such a strong, like, hacker culture and, like, people are just, like, they get excited about things. And so, guy from inference, this one awesome guy from, the infra team are like, “ this is amazing. We’re gonna, like, hack on it.” They build a prototype, it, like, works, and now we- we are just, like, hill climbing on latency and trying to make this as fast as possible, and we wanna, like, launch it in the coming days. so as soon as we hit our, like, latency target, we’ll try to get this out.
Vibhu [00:25:13]: It’s interesting. At the same time of hacker culture, you also, as Sam said, like 99%, one of the most reliable APIs with
Nikunj Handa [00:25:20]: Mm-hmm
Vibhu [00:25:20]: I think probably the most usage, which is your team directly. how should people see decisions API? I feel like a lot of people saw Jev, heard the buzz, haven’t built with it. You’re making it very mainstream.
What Decision Models Are Good For
Nikunj Handa [00:25:32]: Mm-hmm.
Vibhu [00:25:33]: What should people see it as? How should they use it?
Nikunj Handa [00:25:36]: Yeah. I think the main use cases we’ve seen is, like, really fast classification. all the Computer Use demos have been amazing and really cool. I think there will be limitations, of course, in terms of, you know, having Astra, like, write, like, a JavaScript-like script to control your computer, versus having Luna pick, like, one action at a time. I think, it’s not gonna be at the same intelligence level, but, like, maybe there’s some Computer Use tasks that this is good enough for. So excited to see that come through. the other cool prototype I’ve seen internally is people hooking it up with GPT Live. So GPT Live is like, you know, our bidirectional, like, real-time,
Swyx [00:26:14]: Voicing
Nikunj Handa [00:26:14]: A- API. And, it’s built on this, like, model of front-end models and back-end models. So GPT Live is this, like
Swyx [00:26:20]: Think or talker
Nikunj Handa [00:26:21]: Super fast. Yeah, think or, talker thing. So GPT Live is the talker, super fast, really good at delegation, and you have something like Astra sitting at the ba- at the back. But tool calling has always felt, like, really slow in GPT Live. and so people have been, like, putting together these, like, tool calling demos of GPT Live controlling a computer, and it just feels like so much more snappy and natural. So I’m, like, kinda excited to see, like, what people do with Live and with Luna on decisions API. so that’ll be pretty exciting. Yeah.
Swyx [00:26:55]: So I wanna iron this out for people, especially from the product side, because a lot of people have been putting out Jev clones. There’s been about 100 in the last two weeks.
What Makes a Decision Model Different
Nikunj Handa [00:27:01]: Oh, really? That’s amazing.
Vibhu [00:27:03]: The first couple days.
Swyx [00:27:04]: But like, it. Like, they can clone a Jev API, which is honestly structured outputs
Nikunj Handa [00:27:09]: Yeah
Swyx [00:27:09]: Which OpenAI was first to.
Nikunj Handa [00:27:10]: Yeah.
Swyx [00:27:11]: Right? So, like, I think let’s iron out for people what is a decision model, as far as
Nikunj Handa [00:27:16]: Yeah
Swyx [00:27:17]: As far as, like, what is important? It is not just latency. It’s not just structured output, right? Because I could just have Luna as it’- The decision model is priced the same as Luna, right?
Nikunj Handa [00:27:26]: Mm-hmm.
Swyx [00:27:27]: Have turned off reasoning and then have structured output. Do I have a Jev? you know, no, right? And that’s the
Nikunj Handa [00:27:33]: Yeah
Swyx [00:27:33]: That’s the real
Vibhu [00:27:34]: There’s a confidence there.
Swyx [00:27:35]: Yeah.
Nikunj Handa [00:27:36]: Yeah, totally. I think, the way that. So we haven’t trained, like, a new model for this.
Swyx [00:27:40]: Yeah.
Nikunj Handa [00:27:40]: We’re, like, building this purely on top of the same Luna weights that we have.
Swyx [00:27:44]: Oh.
Nikunj Handa [00:27:44]: So yeah. This is, like, really just Luna. And, on top of that, what you’re doing is you’re constraining. So, like, structured output’s a big part of it. you’re really optimizing the inference stack to, like, get very fast on TTFD. And because you can have multiple questions, what you do is, like, you basically run those in parallel,
Swyx [00:28:05]: As a batch.
Nikunj Handa [00:28:06]: Yeah. You run those in the-- as a batch. you-- All sorts of, like, inference techniques people are working on to try to make it as fast as possible. But I’d say, like, at least our implementation of it at the start and this first version is, like, zero-shotting this on top of Luna, to see how it goes. And obviously, you wanna, like, put it out there. Like, this is OpenAI’s, like, classic iterative deployment thing. Put it out there, see what people think, and then, like, we’ll make more model improvements, as needed. so yeah. That’s, the decisions API.
Swyx [00:28:38]: Yeah. And, obviously as a benefit, you have vision. They don’t have vision, right?
Nikunj Handa [00:28:42]: That’s true.
Swyx [00:28:42]: Obviously, Jev’s comes with
Nikunj Handa [00:28:43]: Yeah. Like, we get it for free with Luna. Yeah.
Swyx [00:28:45]: Yeah. I do think that, like, you know, some of the innovations, it sounds like, it’s still to come if it’s still the same Luna weights, which is, like, the confidence stuff, like, the in calibration is something that we’ve talked about on the podcast with, benchmarking calibration. ‘Cause basically, the whole point is that RLHF kind of collapses you towards what you want to hear.
Calibration, Architecture, and the Open Research Questions
Nikunj Handa [00:29:03]: Yeah.
Swyx [00:29:03]: But, like, not actually, like, what the amount of confidence is.
Nikunj Handa [00:29:06]: Yeah. Yeah, totally. I’m eager to see how it pans out. Maybe there’s, like, gonna be. These are gonna be, like, the key areas where we may have to, like, hill climb
Swyx [00:29:15]: Yeah
Nikunj Handa [00:29:15]: With the, with the future model release. But, yeah.
Swyx [00:29:18]: And then architecture-wise, the other thing that’s in the debate, obviously, you-- Nobody knows because Jev doesn’t talk about it, but the two speculations are, one, maybe diffusion model instead of autoregressive.
Nikunj Handa [00:29:28]: Mm-hmm.
Swyx [00:29:29]: But you are able to achieve the parallel, generation in your way. And then the other one is some mech interp type thing
Nikunj Handa [00:29:37]: Mm-hmm
Swyx [00:29:37]: That you’re, like, analyzing the activations and then just outputting
Nikunj Handa [00:29:40]: That would be cool
Swyx [00:29:41]: The weights.
Nikunj Handa [00:29:42]: Yeah.
Swyx [00:29:42]: Which, like, you guys have all done the research on this. People have speculated.
Vibhu [00:29:45]: There have been demos on
Swyx [00:29:46]: Yeah
Vibhu [00:29:46]: Both of these as well. I think Gemini shared a Gemini diffusion, Gemma diffusion on a Jev-style output.
Nikunj Handa [00:29:53]: Oh, sick.
Vibhu [00:29:53]: And, interp people have also, you know, pulled out interp from a middle layer, but this is all speculation.
Swyx [00:29:59]: It’s just like, what are you trying to aim for, right? Because you can achieve the API. Everyone can achieve the API. It’s actually pretty trivial. But, like, then there’s the speed, then there’s the accuracy, then there’s the other calibration features.
Nikunj Handa [00:30:11]: Mm-hmm.
Swyx [00:30:11]: I don’t know what else.
Nikunj Handa [00:30:13]: Yeah. Yeah. No, totally. It’s so cool that this, like, whole space has been kicked off now and people are gonna do so much cool stuff and everyone’s gonna learn from each other. And, yeah, I’m excited about it.
What Developers Should Build Next
Vibhu [00:30:24]: I feel like being on the platform team, a lot of your job is to empower builders.
Nikunj Handa [00:30:27]: Mm-hmm.
Vibhu [00:30:28]: What do you think people should build with decisions API and also Computer Use agents? Any stuff that you’ve- been building with internally that you think really opens up after the new change?
Nikunj Handa [00:30:39]: Yeah. okay, let’s think. decisions API, use cases internally have been pretty obvious. Like, the user ops team was, like, jumping on it. We were like, “We gotta classify all of our support tickets.” what else came up? obviously, there were, like, the really cool GPT Live demos. I’m sure, like, the Codex app team might, like, pick this up and try to do something cool with it. So, you know, like, this whole thing started, like, a week ago, so it’s, like, very early and
Swyx [00:31:06]: Oh, one week.
Nikunj Handa [00:31:07]: We’re excited. Yeah. Yeah, pretty much.
Vibhu [00:31:08]: There was a big push in, evals, LLM as a judge having really low latency there.
Nikunj Handa [00:31:13]: Right. Yeah. That’ll be interesting to see. and then, with the Agents API, we have-- we’re basically, like, having a bunch of first-party products, like, at OpenAI built fully on top of it. we’ve had the Codex security stuff that just went out that’s fully built on top of, the Agents API. We have, sort of the-- we- we are having, like, a meetings type of thing launching today.
Agents API and OpenAI’s First-Party Products
Swyx [00:31:40]: Mm-hmm.
Nikunj Handa [00:31:40]: I think there was, like, a demo. do you remember, like, the plugin extensions when Sam was showing it? There was, like, a demo for, like, you’re in a calendar, you can sort of, like, have your meeting notes
Swyx [00:31:51]: Like, drop into a single
Nikunj Handa [00:31:52]: Flow into like your space
Swyx [00:31:52]: Like, Google Docs type thing.
Nikunj Handa [00:31:53]: Yeah.
Swyx [00:31:54]: Right?
Nikunj Handa [00:31:54]: And so the-- all of that stuff is, like, fully built on top of, the Agents API. and yeah, I’m, like, just excited to see. Like, we’re just getting this out, and let’s see what people build on top of it.
Vibhu [00:32:04]: I think you showed it off very well. The whole edit spaces, pages, collaborate, add in your dot. Like, that’s a lot, so
Nikunj Handa [00:32:12]: Yeah
Vibhu [00:32:12]: There’s a lot of inspiration people can go to.
Nikunj Handa [00:32:14]: Yeah. All possible with Astra, you know. Like, thing- things just move so fast now. Like
Swyx [00:32:19]: Yeah
Nikunj Handa [00:32:19]: People go from idea to execution so quickly, it’s amazing.
Swyx [00:32:23]: Is there something that you want, people to focus on to give you feedback? Like, what-- like, you know, maybe you’re just putting this out there and you want-- and there’s, like, a fork in the road and you want developers to help you decide.
Responses API Performance and Long-Lived Caching
Nikunj Handa [00:32:35]: So I think Agents API and decisions API, they are like, these are our newest products. Would love, like, any and all feedback on that to figure out where to take them. I think, over here, we’re, like, very open on Responses API, which is sort of like our workhorse over here. like, really focused on performance right now, and the performance comes in, like, two main ways. first is just, like, latency. We’ve been, like, rewriting the whole Responses API stack to, like, make it as fast as possible from a TTFT perspective, DVD perspective. So there’s like-- that, like, continues to be, like, a main area of focus for us. The second thing we’ve been trying to do is, like, really go deep on caching, particularly with these, like, personal agents that are, you know, like, basically, like, a single thread that just goes on and on forever. We’ve been, trying to, like, really up our game on caching. We provide now guarantees of, like, cache hits within, like, 30 minutes. We’re actually, like, we-- for one of our users, we just launched, like, a much longer cache window. So we have, like, a 12-hour caching guarantee, that we offer so that you have, like, guaranteed cache hits for
Swyx [00:33:40]: Is that a public API?
Nikunj Handa [00:33:42]: Not yet. That’s in preview.
Nikunj Handa [00:33:43]: We’re gonna, like, try to get that out to everyone as soon as possible. But, like, just pay a little bit more for the cache write, and we, like, guarantee, like, cache reads for, like, a much longer period. So even if, like, your instinct thread, for example, like, you just, like, do something on it and then come back to it, like, three to four hours later, you- you’re still getting the caching performance out of it. And launched
Vibhu [00:34:04]: And you cut the cost there quite a bit too, right, with the new model?
Nikunj Handa [00:34:07]: Oh, yeah. That’s right.
Vibhu [00:34:08]: Like, 25% cheaper, so
Nikunj Handa [00:34:08]: Yeah, with, like, driving down cache reads, yeah.
Cache Pre-Warming and Cost-Efficient Agent Threads
Vibhu [00:34:10]: For builders, they should implement
Nikunj Handa [00:34:13]: Yeah
Vibhu [00:34:13]: Because it’s significantly cheaper.
Nikunj Handa [00:34:14]: Yeah. Yeah. Just, like, building your apps with, like, to be very cache aware and sort of, like, use our prompt diagnostics or cache diagnostics tool to figure out, like, where things are dropping off. And, so the caching part is, like, really important. yeah, I also wanted to talk about pre-warming. We have that in the API now. So, like, if you know that, “Hey, I’m gonna get a cache,” like-- sorry, “I’m gonna get this prompt. I just wanna, like, pre-warm the cache, pay, like, the cache write fee right now, and then, like, have it sort of ready to go for the next 30 minutes for whenever.”
Swyx [00:34:49]: And it can spawn many instances of that thread.
Nikunj Handa [00:34:51]: Exactly, yeah.
Swyx [00:34:52]: Yeah.
Nikunj Handa [00:34:52]: You can just keep going and have
Swyx [00:34:54]: Yeah, just keep messing with the prompt there
Nikunj Handa [00:34:55]: Tons and tons of that. and so, yeah, like, I’m very excited about getting feedback on, like, the low-level performance things that we can keep making Responses API the most performant and reliable way to, like, build on top of an LLM. And then you basically have our, like, new products where I’m just looking for, like, any and all feedback.
Swyx [00:35:15]: Yeah, just use it, right?
Nikunj Handa [00:35:16]: So yeah, just use
Swyx [00:35:16]: Tell us what to
Nikunj Handa [00:35:17]: Yeah. Define our roadmap for us, please. So yeah.
Swyx [00:35:20]: I think for me, the caching thing, great, right? Like, obviously very needed. But at the end of the day, you’re still bumping up against a million-token context
Compaction and Managing Million-Token Contexts
Nikunj Handa [00:35:28]: Mm-hmm
Swyx [00:35:28]: And that’s probably not gonna change for the foreseeable future.
Nikunj Handa [00:35:31]: Mm-hmm.
Swyx [00:35:31]: Like, you still need good compression.
Nikunj Handa [00:35:33]: Yeah.
Swyx [00:35:33]: What is the best practice there?
Nikunj Handa [00:35:34]: Yeah. Yeah, totally. so firstly, OpenAI has its own, like, proprietary compression, comp
Swyx [00:35:40]: Which is in
Nikunj Handa [00:35:41]: Compaction.
Vibhu [00:35:42]: Compaction.
Swyx [00:35:42]: It’s in the agents.
Vibhu [00:35:43]: It’s in the API.
Nikunj Handa [00:35:43]: Yes.
Vibhu [00:35:43]: Agents API.
Nikunj Handa [00:35:44]: Yeah.
Swyx [00:35:44]: You decide for us, right?
Nikunj Handa [00:35:45]: Yeah, exactly. So in the Agents API, it comes built into the harness. and if you’re in Responses API, there’s, like, two ways of doing it. One is what we call server-side compaction, which is you basically tell Responses API that if you ever hit this threshold of tokens, just auto-compact it and, like, go back, or sorry, like, reduce the context, being used. And the second way is, like, /compact, which is, like, if you want full control. So you can, like, /compact at any time
Swyx [00:36:15]: I hear you
Nikunj Handa [00:36:15]: Have your own logic on when to, like
Swyx [00:36:17]: It’s not AGI.
Nikunj Handa [00:36:18]: It.
Swyx [00:36:18]: It’s not AGI.
Nikunj Handa [00:36:19]: Yeah. Yeah.
Swyx [00:36:20]: Yeah. But it, I mean
Nikunj Handa [00:36:20]: Yeah
Swyx [00:36:20]: It is the manual override.
Nikunj Handa [00:36:21]: Yeah, it is the manual way. And like, I don’t know, but a lot of the big coding agents like to do it manually. I mean, like, if you look at the Codex implementation of it in the Code- open source Codex harness, you can see that they use /compact and do it. and, there’s also, like, new, by the way, new compaction techniques that we are working on. Some of them you will be able to see in the Codex harness. Like, it’s already implemented in the Codex harness. And so, they’re like some file-based, systems that we are, like, experimenting with. So yeah, lots of cool stuff going on around in compaction as well.
Swyx [00:36:57]: Cool. we are running out of time.
Nikunj Handa [00:36:59]: Okay.
Swyx [00:36:59]: I think you’ve talked about, a lot about performance and talked a lot about, the new APIs that you’re launching. Can you give us any other hints as to things that you’re interested in as far as the future of the platform is concerned?
Higher-Level Platform Primitives and the AI Cloud
Nikunj Handa [00:37:13]: We’re obviously like very low level. Like, I used to work at Stripe before this, and, at Stripe a lot of the game was like building these higher level primitives and products on top of like the core payments primitives. and, I’m always like curious about what the best way of doing that is in AI. And I think we’ve had a couple of attempts at that. We like had launched assistance API like way back in the day, and like wasn’t really the right fit. We were sort of like going off with this like Agents API, and, it gives you the codex harness, but like where’s like the, what’s the right amount of flexibility to give in that? That’s like an open question. Like how should we like have memory walls and like all of these like higher level like API objects to take away, also like to abstract away more, like storage concepts. Like this is like a whole, like, there’s a whole space that I’m like very curious about figuring out how we design. I think a lot of things in AI are just have a low-level API primitive and see an example harness and go and have your coding agent implement that. But how much of that do we build into the API is like a constant question that I’m thinking about.
Swyx [00:38:24]: Yeah.
Nikunj Handa [00:38:24]: So I don’t know if folks have thoughts on that. If anyone has ideas, it would be super interesting to hear.
Swyx [00:38:30]: Yeah. The analogy I always bring back to, and we’ll end there, is, you’re building an AI cloud, right?
Nikunj Handa [00:38:35]: Mm-hmm.
Swyx [00:38:35]: Like, which is, something that, Sam said a year ago
Nikunj Handa [00:38:38]: Mm-hmm
Swyx [00:38:38]: Where, and you’re, it’s almost like you’re kind of doing the AWS invention and you have to do, okay, this is EC2
Nikunj Handa [00:38:45]: Yeah
Swyx [00:38:45]: And this is S3, and this is like. But you’re doing the AI-native versions of each of these.
Vibhu [00:38:48]: There are a lot of analogies, so you’re pre-warming caches for stuff that you know will be
Nikunj Handa [00:38:53]: Yeah.
Vibhu [00:38:53]: And it’s nice that it’s all exposed to builders
Closing
Nikunj Handa [00:38:56]: Mm-hmm
Vibhu [00:38:56]: ‘cause it just opens up ways that you can build new things.
Nikunj Handa [00:38:59]: Yeah, absolutely.
Swyx [00:39:00]: Okay.
Vibhu [00:39:00]: Awesome. Well
Swyx [00:39:01]: That’s everything.
Nikunj Handa [00:39:01]: Thank you, guys.
Vibhu [00:39:02]: Thank you.
Nikunj Handa [00:39:02]: Yeah.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe- We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks!
In case you’ve been under a rock, here’s a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR:
* June: Launched Claude Tag and Sonnet 5 and Fable 5
* July: Opus 5, /checkup. crossed $65B ARR
* Last month: Fable/Mythos 5.1, and EFS (upcoming pod)
* IPO target $2T, end 2026 ARR estimated $100B
* Cowork/chat merged before did
* Claude Mods
* Dario endorses the same Pacing the Frontier message cosigned by all labs
* Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects
* Today: Sonnet 5.5!
Today’s episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable:
The Future of Mutable Software
Pay special attention to Claude Mods (especially the cheatsheet):
In general this is also the inverse of the other viral tweet from Thariq:
Cloud Brain, Local Hands
And give a try to Claude Projects:
The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we’ll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment.
For those who want Thariq’s writing tips we teased at the start of the pod, watch the full video here:
From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic’s Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.
We go deep on Claude Code’s evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized.
The conversation then turns to agent security and Anthropic’s “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years.
We discuss:
* Why agentic coding went from controversial to the default in less than a year
* Why prompting is still one of the highest-leverage skills for working with Claude Code
* How expert users build a mental model of Claude and what it can reliably one-shot
* Why discovering your “unknown unknowns” matters more as agents become more capable
* Artifacts as persistent, generative interfaces between humans and agents
* How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces
* Claude Tag, Projects, and multiplayer agents and how collaborative agent workflows could evolve
* Why spending more time on the initial prompt can dramatically reduce wasted agent work
* When to use low, medium, high, or max effort for different engineering tasks
* Why frontier models may eventually outperform smaller models on both intelligence and token efficiency
* Why implementation notes can expose decisions the model considered but chose not to make
* Why Claude.md may eventually disappear — and why starting without one can sometimes be better
* Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code
* Model routers, forked agents, and supervisor agents that automatically improve agent workflows
* Why Claude Mods may be an early preview of “mutable software”
* The bitter lesson of harness engineering and why agent architectures go out of date so quickly
* How Claude Tag is becoming an organizational harness for multiplayer work
* Why giving agents access to company data creates an enormous new security surface
* The Exploit-Bench incident where agents discovered ways to communicate and collaborate
* Why agents hacked Hugging Face for scorer code rather than benchmark answers
* How agents chained sandbox and infrastructure vulnerabilities in unexpected ways
* Why increasingly capable agents make traditional security assumptions harder to maintain
* The argument behind Anthropic’s “Pacing the Frontier” proposal
* Why software engineers are increasingly doing two jobs: engineering and keeping up with AI
* Constitutional classifiers, probes, and fallbacks and what interpretability looks like in production
* How Auto Mode checks whether an agent’s actions actually match the user’s permissions
* Why Thariq can see serious AI risks while still having a relatively low p(doom)
Thariq Shihipar
* X: https://x.com/trq212
* LinkedIn: https://www.linkedin.com/in/thariqshihipar
Timestamps
00:00:00 Introduction
00:04:12 Ask User Question and the Future of Agent Interfaces
00:08:29 Artifacts, Projects, and Multiplayer Agents
00:15:37 Prompting as the Core Claude Code Skill
00:21:52 Context, Effort, and Smarter Model Usage
00:28:10 Is Claude.md Going Away?
00:32:49 Claude Mods: Customizing the Claude Code Harness
00:36:35 Model Routing and the Rise of Mutable Software
00:44:40 The Bitter Lesson of Harness Engineering
00:50:49 Claude Tag as an Organizational Harness
00:55:59 Pacing the Frontier and Autonomous Agent Security
00:58:22 Agents Hack Hugging Face for the Scorer
01:05:34 What Happens When Agents Need More Compute?
01:10:32 AI Coding Is Changing Faster Than Engineers Can Keep Up
01:17:17 Probes, Fallbacks, Interpretability, and Auto Mode
01:28:32 AI Risk, p(doom), and Closing Thoughts
Transcript
Introduction: Life at Anthropic and the Pace of Change
Swyx [00:00:00]: We’re here in the studio with our friend Thariq from Anthropic, and I guess generally the Claude Code, I-- there’s, there’s so much, merging of boundaries and you’ve been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you’ve told that story in other podcasts, and you’ve also been talking about seeing like an agent. Most recently you did the top AIE World Tour talk, Field Guide to Fable, which obviously you guys launched Fable, so that was-- that’s cheating. And mostly you most recently also launching Claude Tag, and we’re also gonna be talking about Pacing the Frontier. There’s a lot going on in Anthropic. I guess top of the question is, what’s it like being at Anthropic when there’s so much going on?
Thariq Shihipar [00:00:48]: I think that It is, like. I think you can get whiplash sometimes. I think, like, going. When I joined Anthropic, I joined because of Claude Code. Like Claude Code had just come out and I was like, “This is so good.” And Opus 4 to me was like just, I could not imagine, like, how good it was? And that was, like, a real moment for me. But I was, like, trying to convince, like, my startup friends to use agentic coding, and they’re like, “Oh, no, like, our engineers don’t think it’s good enough,” or something. And I was like, “That’s insane.” and now you, like, fast-forward, 12 months, less, and, like, it’s just like, yeah, the default way that everyone codes, right? And I think that, like, just having to go from, like, selling it to, like, now, teaching people how to be. make the most use of it and be more efficient and things like that is just like a big, like big change. And, yeah, I think, like, it’s just hard to stay on top of everything as a human? Like, I think things happen so fast and like
Swyx [00:01:51]: You just throw more agents at it.
Thariq Shihipar [00:01:52]: Yeah, like that’s like the agentic stuff scales much better than the, like, human stuff where it’s like, oh, like, there are three things happening right now and, like, they’re all emergencies and, like, how do you, like, respond to it? Yeah.
Teaching People to Use Claude Code
Vibhu [00:02:05]: What do you split your time on? You do a lot of technical writing, engineering work.
Thariq Shihipar [00:02:10]: Yeah, so I think that, like, when I joined the Claude Code team, I wanted to teach people how to use Claude Code and I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, I’d, like, do. I was spending some time on the agent SDK first, and I wasn’t exactly sure, like, how the bitter lesson would go, when it comes to, like, harnesses, right? Like, I think sometimes we were like, “Oh, like, what’s after Claude Code?”? And so initially I was like, I just wanna teach people how to use Claude Code and make it easier to use Claude Code. And I think that has just, like, as the harnesses have gotten better and better, that’s like the dominant problem now is, like, how do you use the agents, right? Like, it’s like such a high skill expression thing. So I do that and then I do engineering work. I give talks, but I think, like, when I’m doing engineering work, my goal is to take that feedback that we get from users and also, like, then be able to talk about, like, hey, how to use Claude Code to do engineering. So there’s like a good loop there. Yeah.
Swyx [00:03:07]: Yeah. I’ll-- For listeners, we’ll attach, the talk that you did with Sarah for the Dev Writers, meetup
Thariq Shihipar [00:03:13]: Oh, yeah
Swyx [00:03:13]: Which we talked a little bit about, well, first you do the work and then you talk about the work.
Thariq Shihipar [00:03:16]: Right.
Swyx [00:03:16]: Something like that.
Thariq Shihipar [00:03:17]: Yeah.
Swyx [00:03:17]: It’s sow and reap or
Thariq Shihipar [00:03:19]: Yeah, reap and. Sow and reap.
Swyx [00:03:21]: Something like that. Something like that. Yeah, so, and then just to preview a little bit, we are gonna talk about the evolution of the harness. It has come a long way from just being a CLI. We’re gonna talk about, Claude Mods, which is starting to leak today, because you couldn’t keep it secret.
Thariq Shihipar [00:03:36]: Yeah. yeah.
Swyx [00:03:39]: Yeah, there’s, there’s a lot, there. I think you started off with, like, adding ask user question tool, which people love and hate.
Thariq Shihipar [00:03:48]: Yeah.
Swyx [00:03:48]: Like, I thought it was, like, very innovative, and then now I have, like, my own version. You have your Interview Me version.
Thariq Shihipar [00:03:55]: Yeah.
Swyx [00:03:56]: And, yeah, everyone just has, like, their own stuff. And, like, it no longer matters ‘cause now you’re supposed to, write prompts that create other prompts and loops and all these things.
Ask User Question and Human-Agent Interaction
Thariq Shihipar [00:04:05]: Sure, yeah.
Swyx [00:04:06]: So what’s the state of the art, today? Like, what are people. what are you, like, telling people to do today?
Thariq Shihipar [00:04:12]: Yeah, ask user question was the first time that the model was good at elicitation. I think this was, like, an emergent behavior that I, like, wanted to see if the models could do. I have, like a human-computer interaction background, so I, like, did that in undergrad and grad school. And so this was like. I think it’s like human-agent interaction to me, like, trying to figure out, like, how can the agent communicate with you and extract, the requirements, right? I think that, like, one of the things about, like, that’s difficult as Claude Code has gone broader and broader is that everyone has, like, their own way of using it, and it’s very hard to, like, change the default behavior. So for example, like, if someone asks Claude Code to do something,
Thariq Shihipar [00:04:59]: Sometimes they just want them to do the work, ‘cause they’re, like, maybe a very good prompter, and sometimes they want. like, are not good at prompting? And you need. like, the agent needs to, like, clarify? And so that’s, like, a good split. Like, and the ask you the question tool like, splits along that side where, like, are-- do you feel like you’re good enough to instruct the agent as it is, or is the agent able to, like. does the agent need to, like, pull out more requirements and, like, collaborate with you more and really understand your preferences?
Thariq Shihipar [00:05:27]: I, on the whole, believe that pretty much everyone is more on the latter than the former, that they, like, have more ambiguity and they know less than they want, than they, like, think they know about the problem. but, like, it’s like a interface design problem to make that easy? And so, like, if you’re designing a problem, like, or if you’re going through a problem, like, things like what’s the schema or, like, what’s the call stack and things like that are really important. like, the details in the design are important. Ideally, you want to figure out some of these, like, hard problems ahead of time before starting implementation. And yeah, that’s why they call, like, unknowns, right? And so I think that this will forever be, like, a skill in agentic coding is, like, figuring out your unknowns. So, like, because even if the model is, like, super intelligent- It, like, needs to know what you want? And, like, you have preferences. like, you need to like, pull the, pull that out. and so that’s, like, I think how I’m, what I’m pushing. the question then is, like, how does the agent interact with you? And I think that has been HTML, has been, like, the big way of doing that. And we’ve recently added artifacts, right? And artifacts, I think we’ve done a bad job of, like, or, like, I’ve done a bad job of, like, explaining how to use them fully. We have a lot of property capabilities. They have a database associated with them? And so every artifact can store and write persistent data. They can, like, feed back into Claude? And so, like, one thing that, like, people are not doing yet that I’m trying to, like, encourage is, like, this idea of a dashboard artifact. So you have, like, Claude working on a project long-term. Maybe it’s like a kanban or something. it can store that kanban data in its database. Multiple Claudes can access that data via, like, the artifact MCP, and, like, that artifact can, like, talk to those Claudes as well. And so, like, the. We’re building the primitives for you to be able to have this, like, generative interface via artifacts that will, like, let you surface more of that rich detail from the agents. And I think that, like, almost everything with agents right now is, like, this problem of, like, you think what you want, but you don’t really know what you want, and, like, the agents need a lot of detail, and collaborating with them in the loop is really important. And so artifacts are, like, the, like, way that we’re trying to evolve there. But there’s a lot of work to do because it’s so much more complicated than, like, a multiple-choice question? there’s a lot more, like, detail in terms of, like, diagrams and code snippets and schemas or, like, whatever it is for that problem. But, like, artifacts is, like, the mo-more AGI-pilled way of, like, doing ask user question. So yeah.
Artifacts as the Interface to the Harness
Swyx [00:08:15]: I think one thing that’s unclear to me about these, the artifact stuff is, like, what feedback should go in through the artifact and what feedback should go through a Claude, a chat? Because the more AGI-pilled one is to just feed everything to the Claude.
Thariq Shihipar [00:08:29]: I think the more AGI-pilled one is to go through the artifact. Like, and I think that, like, we imagine in the limit, I think that artifacts will be your interface into the harness? You can, like, comment on this, like, live, like, document of your plan, of the work. you can see maybe, like, multiple agents and different agents are doing this, and that artifact is built for the current work that you’re doing, right? And so, like, each one has, like, slightly different. I think we’re still, like, getting there from, like, an infrastructure perspective. But yeah, I think, like, on-the-fly interface for your harness is probably where things are headed.
Vibhu [00:09:03]: Is there a version of it that’s an abstraction from CLI or chat and you. Because right now, a lot of it is, okay, you’re interfacing with Claude Code, you’re having HTML given back for a mockup. It’s pretty rich. There’s diagrams. Artifacts are ways to connect these together. Why not just do everything that way?
Separating Brain, Hands, and Surface UI
Thariq Shihipar [00:09:22]: Then it becomes, like, separating out, like, where is the inference happening? Where is the intelligence happening? Where is the work happening? like, I think this is like, difference between, like, or, like, some of the distinction between local and cloud, right? And so, I think right now, if you use Claude Code, it’s, like, local and, like, you can spin off remote control, for example, to get some cloud behavior, or you can spin off Claude Code in the cloud, right? We’re moving towards a place where instead of Claudes, like, you message a local Claude, it starts a session locally and it executes, to more like you have a Claude that you message that’s in the cloud that’s running. it can run, like, local, or, like, cloud sessions. This is how Claude Tag works. But, like, over time, we’ll add, like, local hands as well. And so, like, local hands will be the ability for that agent to access your computer if it’s online, and be able to, like, work there. And so it can spin off many different subagents. It can, like, commu- those subagents can communicate with each other, and that’s where the artifact comes in to display all of that work. So you can imagine, like, the. You’re separating out these things. So there’s, like, the surface UI display that’s an artifact and hosted somewhere and has a database and everything. There is the inference intelligence, right, that’s happening on the cloud, and you don’t have to worry about shutting off your computer or whatever, right? and then there’s the, like, hands. Like, and it can be local, it can be in, like, a remote sandbox or wherever you need your work to be done. That’s like unpackaging, like, the Claude Code experience right now where, like, right now it all happens in one place, right? So.
Multiplayer Agents, Claude Tag, and Projects
Vibhu [00:11:00]: How do you see, like, the multiplayer side of that? So say teams want to work in this way. Right now it’s very individual, but how do you see the future of multiplayer? Like, right now, I guess there’s Claude Tag, which is a version, but.
Thariq Shihipar [00:11:12]: We’re launching projects. And so projects is the, like, this abstraction that’s like Claude Tag, but on our Claude products, right? So you can message it and, like, it will do the Claude Tag-like stuff, like spinning off subagents. So We think with multiplayer. Like, Claude Tag is, like, a little bit more native multiplayer because it’s just, like, in your Slack and the permissions are all figured out and stuff like that. But I do think multiplayer is, like, an important part of the story and, like, that will need to get tied together more. Like, you can imagine how complicated it gets when you’re like, oh, you have hands, but now you have other hands in other people’s computers too, and, like, you need to, like, permission them or, like, you have, like, your MCP and someone else’s MCP, and how do you figure out how to use them, right? It gets, like, quite complicated. And Claude Tag does a good job of, like, sanding down all of these issues, right? So that, like, when you have, yeah, Google Docs, how does it access Google Docs, right? Like, it accesses through the shared Claude MCP, or it can access through your local credentials as well if it doesn’t have access. But yeah, I think Claude Tag is our multiplayer, product, and it’s really useful for these, like, things that are inherently multiplayer. Like, okay, like on-call, for example, incidents are inherently multiplayer. You want to tag Claude, you want multiple people to log in, you want it to be able to find context. I think whenever I’m, like, working on something and I want, like, privacy or security or, like, I want other people to review it’s really nice to, like. I’ll have a channel per project and I’ll, like, at legal, for example, be like, “Hey, like, I want to ship this. Can you, like.” Like, here’s. Like Claude knows everything, just chat with it. And that way legal gets precise answers, on like what exactly is shipping into the code, and I don’t need to be in the loop, right? So I think like multiplayer is getting like more and more like, yeah, everyone can participate with Claude. I think Claude Tag is like that product and like projects will start off single player and will like, expand.
Swyx [00:13:14]: I think there’s a question about like maybe dual questions about identity and the unit of isolation.
Identity, Permissions, and Isolation
Thariq Shihipar [00:13:20]: Yeah.
Swyx [00:13:20]: Claude Tag, you specifically chose to make it its own identity
Thariq Shihipar [00:13:26]: Yes.
Swyx [00:13:26]: Which is like, a controversial choice. There’s, there’s other ways to do it.
Thariq Shihipar [00:13:30]: Yeah.
Swyx [00:13:30]: Claude Projects probably it sounds like, if it’s anything like ChatGPT Projects, it is, the isolation is that artifacts, that cloud instance, everyone’s collaborating on this. It’ll. It sounds like, it should be like if you’re, if you’re collaborating with legal on a thing, like that channel should be a project, right? Like it’s not yet
Thariq Shihipar [00:13:50]: Yes.
Swyx [00:13:50]: But it. that’s the natural next step.
Thariq Shihipar [00:13:53]: Yeah, like I think in Claude Tag, it’s effectively. Like Claude Tag, you have to do your own arrangement. And so Claude Tag, yeah, each channel is like you can name it as you want, and I name
Swyx [00:14:04]: Yeah.
Thariq Shihipar [00:14:04]: Like each feature
Swyx [00:14:06]: Yeah.
Thariq Shihipar [00:14:07]: As a channel.
Swyx [00:14:07]: And, but I think like there is some trans- like it’s unclear when there is transference, because let’s say it is. if you have a coworker
Thariq Shihipar [00:14:14]: Yeah.
Swyx [00:14:14]: Who is tagging on all these things, yes, there is transfer
Thariq Shihipar [00:14:16]: Yeah.
Swyx [00:14:16]: Because it’s the same person. but with Claude, it’s unclear if it’s like necessarily like, well, no, you don’t know any of. you don’t know about the other stuff. You should only use this stuff.
Thariq Shihipar [00:14:25]: It’s like the tip of the iceberg meme, right, where you can like. This is what we spend so much time on
Swyx [00:14:31]: Yeah.
Thariq Shihipar [00:14:31]: Is like there is like infinite surface area of like, okay, you want Claudes to. Not infinite, but like there’s like surface area, a lot of like, surface area to figure out of like permissions and visibility and like how can you let Claude operate as well as you can, as safely as you can? And obviously, this is very important to us because like security for our code base is very important. And so we’ve put a lot of time into this. Yeah, there’s so many like edge cases you can figure out where it’s like, oh, like, yeah, this Claude in this channel has different permissions, but it can message another channel, and can’t it exfiltrate data that way? Or like can you like. What if it uses your MCP and then messages someone else? Like there’s like so much, and we’ve like really put a lot of work into sanding it down.
Swyx [00:15:14]: Yeah. Lots of work. okay. Fable?
Fable and the Meta-Skill of Prompting
Vibhu [00:15:18]: Fable, you wrote two good articles. you’ve written many good articles
Thariq Shihipar [00:15:22]: Yeah.
Vibhu [00:15:22]: But on, Field Guide to Fable, Building Claude Code. I’m curious from what you’ve seen, is there any common patterns that you see in like top users at Anthropic externally? Like what are best practices for getting the most out of Claude Code?
Thariq Shihipar [00:15:37]: The like meta skill I say is like prompting is like very important? And like that. Like I think this is like not trivial to say because I think a lot of people are like, “Oh, prompting doesn’t matter. It’s just like I can just say a sentence and Claude will do it.” And I think prompting is really this like, this. It’s like public speaking, like, or writing or something, and for a specific audience, and that audience is Claude. And you need to like build a mental model of Claude and how it thinks and how it works, right? And so that’s like the most important skill in working with Claude Code is like having this mental model, right, of Claude and like what it can do well, what it can one-shot, what it can’t. And so many people when you see prompting, they’re just like, they’re short prompts, but they have such a good mental model of Claude and of like the code base and things like that like it’s effortless? But it’s like high skill ceiling. So like that work of like, spending a lot of time prompting and building mental models of how, and intuition for how the agents work is really important. And then I think like the next thing is like the unknown stuff we talked about earlier, where it’s like being able to find out like your, what you don’t know or what you haven’t written down, learning about like different things. I think as Claude can do more and more things, the likelihood of you doing something out of distribution for you and like you have low domain knowledge on is very high? And the more you can like learn the vocabulary to be able to prompt Claude, it becomes really important. And so like I think the most important unknowns are the unknown unknowns, where you’re like, I just like don’t even know that this exists, right? Yeah, exactly. I think that’s like a illustration of like the map and the territory, right, where you’re like, “Okay, this is my prompt,” and the territory is like the actual like work that the agent needs to do, right? And if you are like very precise, you can give more precise things, right? So like for example, in design, I’m not very precise. I’m not a designer, so I say like, “Give me like eight different mock-ups.” But if I was a designer, maybe I’d be like, “Oh, hey, here are some reference sites.” Like, “I want this type of font and this type of like look to it, and here’s like a few different components to like visualize. Here’s a Figma MC board to bring in,” like. And so you can just be so much more precise with that language. And if you’re not a designer, you just need to like try and learn the language or learn the unknown unknowns. And this is true of like everything, I think. Like the more, like you can work with Claude to learn like how things work, the better your prompting will be. I think another good example of this is like game design, like where a lot of people are like, “Oh, like I can vibe code a game now.” And they’re like, “It’s not fun.” And like it’s just like the thing about game design is like every one of these choices has like a lot of
Taste, Domain Knowledge, and Learning the Vocabulary
Swyx [00:18:25]: Variations.
Thariq Shihipar [00:18:25]: A lot of like craft to them. So it’s like, oh, okay, like when you’re making a flying game, the feel of the plane and the like, way it responds to your controls has a lot of like. Like, a game designer would spend like days on that. Do? and like
Swyx [00:18:44]: To me, that’s what taste is, right?
Swyx [00:18:45]: Like it is like from the possible space of one thousand mathematically valid answers
Thariq Shihipar [00:18:49]: Yeah.
Swyx [00:18:49]: Here’s the one that is the humans will like.
Thariq Shihipar [00:18:51]: Yes. Yeah.
Thariq Shihipar [00:18:52]: I think with taste, I’m like torn on this word ‘cause I think you’re right, but everyone has different definitions, and it sounds kind, sounds like low skill or like elitist almost, where you’re like, oh, like there are certain people with taste?
Swyx [00:19:06]: It’s like taste is what I call taste.
Thariq Shihipar [00:19:07]: Yeah, exactly.
Swyx [00:19:08]: And it’s like these guys don’t have taste.
Thariq Shihipar [00:19:09]: Yeah, exactly. Oh, like an engineer doesn’t have taste. Like I, the like founder, have taste.
Thariq Shihipar [00:19:14]: ? And I think that’s not true. Like I think like the engineers have a lot of taste for these particular like problems? And I think everyone has taste for particular problems. I think like Jason Liu, like say like in order to, yeah, have taste, you have to eat?
Thariq Shihipar [00:19:32]: And I really like that, where it’s like, okay, you have to like do a lot of things. You have to like iterate and figure out what you want, what you like, and, like build that like domain
Swyx [00:19:41]: Yes
Thariq Shihipar [00:19:41]: Domain vocabulary. And then when you’re prompting, you’re like synthesizing all of that for a product.
Swyx [00:19:46]: Isn’t it annoying when someone else says it better than you?
Swyx [00:19:48]: It’s just like, f**k, I have to quote this guy forever.
Vibhu [00:19:51]: Having to quote Jason Liu forever.
Vibhu [00:19:53]: He’s gonna love this.
Thariq Shihipar [00:19:55]: So I get prompts, more than that.
Vibhu [00:19:57]: And sometimes it’s not even that. Sometimes it’s just intuitive, right? Like you don’t realize you even want something till a model puts it out, and you’re like, “Oh, this just feels immediately better,” right?
Voice Prompting and Information Density
Thariq Shihipar [00:20:07]: Yeah, exactly.
Swyx [00:20:09]: One thing I go back and forth on is I feel like the way I prompt half the time, let’s say I use voice.
Swyx [00:20:16]: Did I say voice? Other people have voice. that is the opposite. That is just like me rambling for like two minutes Pressing down the function key and then let go, and then like hopefully it figures it out. And oftentimes it does.
Thariq Shihipar [00:20:26]: Yeah.
Swyx [00:20:26]: But it’s not as thoughtful as like a structured prompt with like Well-run communication as though it’s a PRD or a memo. Is that in line with how people do this? There’s like bimodal prompting where there’s some prompts where you spend a lot of time upfront and other prompts you just dash it off?
Thariq Shihipar [00:20:43]: I don’t think the voice is necessarily low. Like I think it’s like more like how much information is in the prompt. like the model can. Like you can and like add some sentences
Swyx [00:20:53]: Right
Thariq Shihipar [00:20:53]: And be like, “Oh, like I changed my mind,” like in the middle of the prompt, and it will be able to follow that perfectly? So I think the like actual format of the text is less important, but then like the ability to. Like how much information is in it, right? And I think for voice, a lot of times, going back to like human-agent interaction and like for a lot of people, it’s just way easier to talk than to like type? and I. If that gets more information out of you, like that’s better.
Vibhu [00:21:21]: At some level, it feels like just giving the model as much context
Thariq Shihipar [00:21:24]: Yes
Vibhu [00:21:24]: Over prompting before you kick off is a best practice. I don’t know. A lot of the times, like when I was first trying out Fable, I spend a solid 30 minutes like really crafting a long prompt. This, I think, is a response of models running for longer and longer, right? It’s still a little difficult to nudge them as they’re in like, in the loop, but I just like intuitively spend more time kicking off that first prompt and working with it a lot.
Spend More Upfront, Iterate Less
Thariq Shihipar [00:21:52]: My personal opinion is that if I was a software engineer, if I was like, just running my own startup, for example, I think I would mostly fit, stick to a max 20x? like maybe verification and so code review are like separate things. But I think like what I see a lot of times is people hit rate limits when they’re doing this like, oh, like it did a lot of work and you’re like, “Oh, I don’t like this.” Like, “Can you like undo this and redo it?” And then you’re like iterating on this like thing that the model could have done if you had like spent more upfront time or given it better context? And instead it’s like you’re like, “Nope, don’t like that design. Try this.” Or like, “You messed this up,” or something like that. And then that just eats up so much more of like, your usage. And so that’s like, I think maybe like a key like tip both for like efficiency as well, right? And yeah, I think like context, and not just like context on like what the goal is good, right? Like are you building a prototype or is it like a production thing? Like where can you spend compute or when, where can you not spend compute? Like I think you have to give the model permission or like not permission to do things sometimes where, like it doesn’t know intuitively how much you want to spend on this task, right? And you can use effort for this. So I did-- I’m working on a blog post about that where it’s like, if you want. For like we see that effort scales with the complexity of the task. So for security, effort gets like way more results. Like high effort versus like low effort gets, like changes the evals a lot. But for software engineering, it doesn’t change it a huge amount because effort is mostly spent on the verification and the like edge case testing and things like that. And so like being able to like give the model that guidance of like, “Hey, this problem is something that I think I want you to spend a lot of time verifying and edge case testing,”?
Effort, Model Choice, and Verification
Vibhu [00:23:43]: How about model in the mix? So, there’s Opus and Fable with effort.
Thariq Shihipar [00:23:47]: Yeah.
Vibhu [00:23:48]: There’s also Haiku in there.
Thariq Shihipar [00:23:49]: Yeah. It’s not quite true yet, but it’s very close where I think the frontier models will be Pareto dominant over like almost everything. like maybe. And sometimes I think Opus might be Pareto dominant. Do? Like I think depending on like how things, like shake out if it’s like a newer version of Opus. But I think that like increasingly it’s just going to be like the smart model is going to be able to like do the simple task for less tokens than the like the other models because of verification. With verification, in the limit, your model doesn’t need to verify, right? If it’s a perfect model, it just does the work once and it’s like, okay, like you, I did it? And increasingly with Fable, I’m like, I’m like, “Dude, you don’t need to spin up Chromium and screenshot all of these things.” Like I see it. Like you did it, right? And so a lot of the. At higher effort, you spend more of those tokens verifying. But if you’re working on simpler problems, and a lot of software engineering is like well, like in Fable, like low and medium stability, it can spend less tokens verifying. And as the models get smarter and smarter, they will just be able to like, “All right, done.”? Like, I can run the lint for sanity’s sake, but, like, I, like, know it lints? Like, you don’t even need to do that. And that will be so much more token efficient than, like, the smaller models. Yeah.
Swyx [00:25:15]: Is there a good, practice on our side that we can use to see if we’re using too much effort? Like, I freaking
Thariq Shihipar [00:25:23]: Yeah
Swyx [00:25:23]: Hate wasting time on that stuff.
Thariq Shihipar [00:25:24]: Yeah. I know what you mean. I think, like, so in this blog post, my rough distribution is, like, code review and security should be, like, high or max and, like, software engineering
Swyx [00:25:37]: You said recommend mix settings per domain.
Thariq Shihipar [00:25:37]: Yeah. I think, like, if you’re doing, like, UI or something like that, like low and medium, I think is you’re building, like, an API and you want to make sure, like, you cover enough edge cases? And so I think building, like I said, that mental model of, like, how things work across these distributions is, like, yeah, part of the job.
Implementation Notes and Decision Logs
Vibhu [00:25:56]: This is more intuition-driven or eval? Because I’m guessing this would change as you go.
Swyx [00:26:00]: He has evals.
Thariq Shihipar [00:26:01]: Yeah. So what I did in the blog post is I go over all of the terminal bench evals. So there are, like, 70 problems and I’m show that, like, okay, like, in the security problems it does more. and then I also, like, look at some of the transcripts just in terms of, like, how-- what does it answer, what does it forget or something. And a lot of times, this is another prompting tip I have, is, like, asking it to make decision notes or implementation notes because, in every eval problem that it faces, it thinks about the correct solution, and decides not to do it. it’s like, oh, like, here is the answer. What if I did this? And then it’s like, oh, probably not? and then keeps going. And this is, like, the majority of the failures, at, like, a higher max level. It’s very rare that the model just doesn’t know how to do something. If you just have these implementation notes, then you can review and you can be like, “Oh, I want you to do this thing that you didn’t do.” The models are getting better at surfacing that overall. Like, I see in the transcripts of Fable 5.1, like, when it does this output, it will call out its decision-making as well. but making this more explicit in the harness is better. And now we’re, allowing ways of you modifying the harness so you can, like, add some
Vibhu [00:27:23]: Ooh.
Thariq Shihipar [00:27:24]: Calculate with there. Yeah.
Swyx [00:27:25]: Yeah. So I do wanna call out two things that you mentioned that I think exist outside of prompting. One is like, let’s, let’s call it the prompt that is so important that it shouldn’t be in a prompt. It is in Claude.md or Agents.md
Thariq Shihipar [00:27:38]: Yeah
Swyx [00:27:38]: Which is like goals, right? Like your situation, your goals, the things that you want, the thing. and then second of all is the decision log or the experiment log or whatever log of traces that you might want to survive the current session to do those things. Those are, like, externalities that there’s no standard. There’s no-- It’s not like skills. It’s not like MCP. There’s no standard. It’s, it’s just like it’s a markdown file. first of all, is that right? Is Claude.md going away? You have a documented dislike of, Agents.md, but you’re gonna do it?
Claude.md, Agents.md, and Model-Specific Instructions
Thariq Shihipar [00:28:10]: Yeah. Okay. So Agents.md, yeah, like, we’re, we’re gonna do it. I think it’s just, like, different models are very different from each other? But I realize that it’s, like, such a pain to, like, maintain different ones? And yeah, like, as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit, Claude.md goes away, and maybe not even, like, that far. Like, I think, like, I think that right now it might be better to start a new project without a Claude.md.
Swyx [00:28:44]: Yes.
Thariq Shihipar [00:28:44]: I think that, like, maybe if you see very repeated failure modes, you add them to your Claude.md. The really tough thing is that this changes per model. And so, like, if you’ve added a bunch of failure modes or, like even
Swyx [00:28:57]: So you need Fable MD, you need Opus MD.
Thariq Shihipar [00:28:59]: Or well, even Fable 5.1 versus Fable 5.
Swyx [00:29:03]: Yeah.
Thariq Shihipar [00:29:03]: Like, it is annoying. Like, I’m not like,
Swyx [00:29:05]: Yeah
Thariq Shihipar [00:29:05]: Like, we don’t, like, do this on purpose? It’s just, like, how the models work, right? And so, like, maybe, like, Fable 5 had this, like, failure mode that Fable 5.1 doesn’t. And if you keep this context, this running log of a bunch of different failure modes, they will probably over constrain Claude? And so this is like. we just added evals plugins for skills.
Swyx [00:29:28]: Yeah.
Thariq Shihipar [00:29:29]: And so now you can eval if a skill is better. I think Daisy on our team did this. And so, yeah, this is like we’re trying to work on this. We know it’s, like, you still have to spend tokens on it and, like, it’s not, it’s not perfect, but it’s, like, we’re trying to help out with this problem.
Swyx [00:29:44]: And so, and as far as prompting goes, the one tip I wanna offer is, something I have told people a lot is sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication. So I’ve referred to-- This is an executive comms workshop from Heavybit that is the best I’ve ever seen in my career. And they teach this thing called the SCQA model. Just Google it. It’s a, it’s a thing. Like, people have done prompting for decades. It’s just called executive communication. It’s like when one person has to communicate to thousands of people down the org chart, this is what you do. so situation, complication, question and answer, is how you write the memo. but obviously sometimes you don’t have the answer, but you can at least list out the SC and Q, and then they have some examples in there. So just leaving breadcrumbs for people if they want to explore.
Underrated Prompting Patterns and ELI5
Vibhu [00:30:31]: Before we move on, I wanna ask you, any other underrated tips, ways people could get a lot of value from Claude Code that they’re not using?
Thariq Shihipar [00:30:41]: Yeah, I think a lot of them are in the, this unknowns, like, doc. Like, I give a bunch of example prompts, like, using it for brainstorming, using it to quiz you after. we added this, like, explain it like I’m five skill which is a very short prompt. And it doesn’t even say explain it like I’m five. It’s like the key word of this prompt is big pictures, few words. like, that’s like the main thing. And it is shockingly good? Like, you, like, I think I tweeted about this and it’s like /eli5, and, like, you can install it as a plug-in. But yeah, it’s, like, way better at just cutting through the BS and being like, yeah, exactly right here. So the diagrams are, like, quite clear. I think one of the things that is true with artifacts is, like, they put too much text in and people are not reading the artifacts? And so, like, this simplifies it a lot more. And, yeah, this came out of, like, just people at Anthropic, like, going through very complicated incidents and being like, “What is happening?”? So, this one I think is great, yeah.
Swyx [00:31:47]: My version of this is the, it’s like test your understanding. Give you a few choices and then, like, if you get it wrong, you have a mismatch between what you think is happening versus what’s happening.
Thariq Shihipar [00:31:58]: Yeah. I think this is one of those things that everyone loves talking about, and then very few people really do. Like, I think
Swyx [00:32:05]: Really helpful.
Thariq Shihipar [00:32:07]: Yeah. But most people just don’t want to get quizzed about something? Unfortunately, I think this is one of the, like, things that we need to, like.
Swyx [00:32:16]: What’s the opposite of ask you the question or ask you the question before the thing?
Thariq Shihipar [00:32:19]: Yeah.
Swyx [00:32:19]: This is after the thing.
Thariq Shihipar [00:32:20]: Exactly. Yeah.
Vibhu [00:32:21]: It’s a good way to stay grounded of, like, do you even know what you’re doing, right? The worst case is when people send you slop and they haven’t understood what they’re asking for or what the output is, and it’s like, “Dude, I don’t wanna read this. Do you even know what it is?” So, you make it a rule for yourself that before you send stuff, you should at least know what’s implemented.
Claude Mods: Customizing the Harness
Thariq Shihipar [00:32:41]: Yes, but so you could make this a mod and you could build your own mod to, like, make sure you test it. So yeah, you can do that.
Swyx [00:32:49]: All right. Let’s get right into it. What is Claude Mod, and what is this diagram showing?
Thariq Shihipar [00:32:54]: Yeah. Okay, so Claude Mods is you can customize the entire Claude Code harness, and we’re going to. If you have requests, we will, like, let you, like, please let us know. We’ll add more and more. This works for CLI, it works for desktop. maybe it will work for Claude Tag in the future. I don’t know. Like, we’re trying to make this very extensible. You can see this reference sheet. I don’t want people to get overwhelmed by it? At a high level, you can customize both the execution of the harness, and the UI of the harness. And so, like, you say on that Tetris example from Boris, that’s like customizing the UI, right? Like showing, like, Tetris in the game.
Thariq Shihipar [00:33:35]: But, like, let’s say that you wanted to do this thing where you had. you tested your assumptions or, like, tested your understanding after every project, right? What you would do is you would ask Claude to make this plug-in. It would spin a classifier after every prompt. And so, like, at the end of each turn, you would spin off a sub-agent or, like, a forked agent. A forked agent is, like, maintains the prompt cache, right? So it’s like a, like one of those unintuitive things where you can fork and do, like, a little request, and it’ll be very cheap because the entire prompt cache is, like, done. And so you can be like, “Has this task been completed?” like
Swyx [00:34:18]: This is how you do BTW and all those.
Thariq Shihipar [00:34:20]: Yeah. The underlying forked agent, yes. But so you can, in the f-fork sub-agent, you can say, like, “Has this task been completed? If so, return true.” And then in your hook, or in your, like, plug-in mod, or sorry, like, in the sub-agent probably, you would say, like, “If true, give me a quiz.” give me questions and answers, and then, like, in a JSON format, and then you’d parse it, and then you display above the prompt input, this list of questions, right? And so this is something that’s, like, slightly token-intensive because, like, you have to do it after every end of the assistant turn. But it’s, like, a lightweight classification, and then you can, like, get this quiz, and then you’ll see, like, Claude will always do it for you. You don’t need to remember to do it. There are lots of these, like, tips that we’ve talked about, right, where it’s like, oh, implementation notes. You can also add a tool for implementation notes now. And so, like, this tool that I’m adding is, like, register, like, I think assumption is what I’m calling it, but, like, maybe I’ll change it around. And this is a mod. And so, like, you give it a register assumption tool, and then it will keep a list. It’ll. Every time it does it’ll keep a, like, add to the list, and then at the end it will display those assumptions? Another mod I’m working on is a model router. And so, like, internal, like, Claude model routing, right? So it’s. This is, I want to say the reason we don’t do model routing by default is, like, it’s a hard problem? And like
Forked Agents, Assumption Tracking, and Model Routing
Swyx [00:35:51]: You will get it wrong.
Thariq Shihipar [00:35:52]: Yeah, you, like, yeah, you will, like, accidentally use, like, Fable for a hard problem or Sonnet for
Swyx [00:35:57]: Yeah, if you have auto approve, but you don’t have auto mode.
Thariq Shihipar [00:36:01]: Well, you will have auto. Like, you don’t have, like, auto routing or something.
Vibhu [00:36:04]: You don’t have auto mode for model picker.
Thariq Shihipar [00:36:06]: Yeah, exactly. So
Vibhu [00:36:07]: I’m getting the rough question of, like, how much do you open this up and how much do people have to think about this? Like, when you talk about prompt caching and building a router, it seems like you could easily build a mod that routes per query, and I’m just killing my plan very fast, right? I guess my question is more so, like, what is, like, a product talk like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to go through
Swyx [00:36:33]: Oh, definitely power users, right?
Thariq Shihipar [00:36:35]: Yeah, I think it is power users, but, like, the nature of Claude Code is that so many people are power users? Because it’s easy to share things, like you can. Like, one person can make a good model router thing that doesn’t break prompt cache all the time, and then you can, like, compose them. Another cool thing about the plug-ins is that they can hook into and compose with each other. And so I have, like, a mod that will, like, create a mode selector at the top, and any plug-ins can register to be a mode. And so, like, the auto router can be a mode, right? Or, like, you can have a mode that’s, like, artifact mode, where it’s like it primarily talks to you in artifacts. like, you can toggle between plan mode? And so, like, you can create more and more of these modes. But the ability to create modes is in it itself a mod? And so there’s a lot of richness here, but we do want to make it fairly easy. We want to be-- make it so that you can just, like, install someone else’s. You can ta-- you can chat with Claude and, we’ll, like, make sure that it understands the nuances of things like prompt caching and stuff, so it can, like, warn you. This is, like, not extremely complicated behavior for Claude, I think, but we should have just a good skill on how to make mods. and yeah, we’ll see how we go. But I do think that this is, like, a preview of, like, mutable software, and, like, how, like, generative software, just like you can customize safely. If enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this?
Power Users, Modes, and Mutable Software
Swyx [00:38:13]: And by the way, you, we have, you have another cool tweet about how, there’s the infinite money button, which is like make your SaaS, consumable by agents. I think mutable software is interesting and, other people have also tried to do it. I think the hurdle comes when you can do everything, then people, users get, tend to get confused. So usually the stuff that works is just like one opinionated flow. This is in the side of less opinionation. It’s just like, well, more power to power users. And I think probably unlocked by AI, where, like, you can just prompt for whatever the thing is.
Thariq Shihipar [00:38:47]: Yeah, or there can be a skill that gives the opinions?
Mods vs. Hooks vs. Artifacts
Swyx [00:38:50]: Yeah.
Thariq Shihipar [00:38:50]: And then, yeah.
Swyx [00:38:51]: So knowing a little bit about, like, TypeScript and build systems and all these things, the closest-- I’m very curious that the team who worked on this, if, I don’t know how close you were to them, if they drew any inspiration from build systems like Babel, Webpack, all these, like, old school things. Because it sounds very similar, like the plug-in ecosystem of those things where they can compose with each other.
Thariq Shihipar [00:39:11]: Yeah, I’m not deep in the technical details, but I do know it was a collaboration with someone on the Bun team and someone on the Claude Code team.
Swyx [00:39:17]: Yeah, it’s a build system mecca.
Thariq Shihipar [00:39:19]: Yeah. Exactly. It’s, it’s very exciting. But yeah, like, agents can just do this very complicated like, extensibility into your software now. And so, yeah, like, another reason to, like. If you run a startup, like, you can just prompt Claude and be like, “Hey, like, could we make an extension system? Like, what would that look like?”?
Swyx [00:39:37]: Yeah.
Swyx [00:39:38]: And I just really wonder, like, you had hooks in the past and plug-ins, all these things. So what specifically will mods be able to do that those things could not do?
Thariq Shihipar [00:39:47]: Internally, we were originally calling this function hooks. And so, like, that’s, like, gives you a little bit of an idea where, like, hooks register a, like an event to happen and then, like, a script to call. And this inside of the, like, TypeScript runtime is running things. And so, like, you get some benefits of just, like, it has a bunch of things in the Scope with, like, for example, like how many turns is in this conversation, right? Like, how many tokens have been used? Like, et cetera. Like, what are the messages? Things like that. So it has a bunch of messages that can be used. And then it’s just, like, a lot more hooks. So we have, like, or a lot of, lot more, like, things you can register on. And then you can do because of the. because it’s all happening in process, you can, spawn sub-agents, with four contests and contexts and stuff. And, like, that will return. You can parse the results of those. You can use structured output to like, return them. and then you can modify the UI, which you can never do in hooks. So, yeah.
Swyx [00:40:50]: Yeah. Yeah. So modify UI, this is why you showed the Tetris example. Does it also ex-extend to artifacts? I assume it does.
Thariq Shihipar [00:40:57]: You-- Like, artifacts are like a different way of customizing it. like, you can definitely. One of the mods I’m working on is, like, this dashboard mod, which will, like, prompt Claude to maintain a dashboard, that’s an artifact. But they’re like, slightly orthogonal, or not orthogonal. They compose with each other in different ways. Like, mods are, like, a little bit more, like, in your Claude Code harness, changing the agent loop? And, like, the UI is, like, an added benefit. and then artifacts are just like you want to, see things at a high level, very inter- highly interactive. like, the affordances can be a lot bigger than, like a TUI or even in our desktop.
Next Steps, Supervisors, and Persistent Guidance
Vibhu [00:41:40]: I’m guessing you’ll have a good blog post on the differences, because right now you can also, make a loop that outputs to an artifact that’s an interactive dashboard, but you can also do it with a mod. There’s just some thinking about making a hacking on a harness when we don’t know much about the harness, right?
Thariq Shihipar [00:42:00]: Well, something I’m excited about with mods is, like, there’s so much things with Claude Code that you just have to remember? You’re like, “Oh, like, let me do this, and then let me call the dashboard skill that does the loop,” and things like that. And, or like, “Let me test my assumptions afterwards.” And I think, like, if you do all of these things using these little classifiers and stuff, and you’re like, “These are the things I care about. This is what I want to do,” you can, like. You don’t have to remember as much. One more, like, mod I’m working on is a next steps mod that
Swyx [00:42:28]: I have-- I was gonna say, I have a next step skill. I always run next steps.
Thariq Shihipar [00:42:32]: And does it have access to your skills? Like, this is one of those things where I’m like.
Swyx [00:42:37]: I think so.
Thariq Shihipar [00:42:38]: Okay. Yeah, probably
Vibhu [00:42:39]: Do skills need specific access to
Thariq Shihipar [00:42:41]: Well, I think there’s
Swyx [00:42:41]: Don’t they always have
Thariq Shihipar [00:42:42]: I think there’s, like, specific prompting, I guess, to, like, know your skills. Like I think Claude forgets them sometimes throughout, like, the thing. But anyways, the idea of, like, yeah, next steps that also are like, “Oh, hey, this has happened. Use the explain skill to explain to you what happened because this seems, like, quite complex,”? Or, like, yeah, “Use your unknown skill. It looks like you are, like, asking the model to, like, iterate on these small changes. It seems like you could prompt better.” like, “What if you did this?” Right? So, I think, yeah, like spending more compute there. Yeah.
Swyx [00:43:20]: And it should always come out as multiple choice. we have, I have
Vibhu [00:43:23]: We have his skill.
Swyx [00:43:24]: My next step skill is like this.
Thariq Shihipar [00:43:26]: Okay, perfect. Yeah.
Swyx [00:43:27]: You can steal it.
Thariq Shihipar [00:43:28]: Yeah.
Swyx [00:43:29]: Like, but like, for me, it’s all-- I think models really always need to be reminded, what are you trying to do here?
Thariq Shihipar [00:43:35]: Yeah.
Swyx [00:43:35]: Look at the whole transcript and go like, oh, was this original goal? Did your solution solve it? Were you lazy? If you’re lazy, maybe there’s a reason. Maybe you needed approval from me. Maybe you needed, there’s two things you wanna suggest. So it’s, it’s a little bit like the modification of the ask user question or interview me skill. so it’s next steps.
Thariq Shihipar [00:43:55]: Yeah, exactly. And again, the benefit of doing it with mods is you can do it as a fork sub-agent, and so it doesn’t remain in the context afterwards. So you have this, like, idea of like, okay, the model is doing its execution and you have this almost like supervisor, like, that is like making sure that you can do like the next steps well. So yeah.
Swyx [00:44:15]: Yes. I do have two panels and like I often try to have a supervisor thing, keep the high-level context and then the implementation
Thariq Shihipar [00:44:21]: Yeah
Swyx [00:44:22]: Detail in another agent.
Vibhu [00:44:23]: I feel like a lot of this abstracts away as models change? The, like, half an hour ago you said bitter lesson of harness engineering
The Bitter Lesson of Harness Engineering
Thariq Shihipar [00:44:31]: Yeah
Vibhu [00:44:31]: And we’re on the other extreme right now, I feel.
Swyx [00:44:33]: Well, so yeah, exactly. If everything’s customizable, what is Claude Code, right?
Thariq Shihipar [00:44:37]: Yeah.
Swyx [00:44:37]: And which I talked to you about last night.
Thariq Shihipar [00:44:40]: Yeah, I think that this is. I think the bitter lesson is unintuitive? In terms of like. Also, like we’re misusing a little bit of the bitter lesson here where it’s like, it’s more about like scaling and compute and stuff. But like, I think there is something where it’s just like. I think I use it as an approximation here to say that harnesses go out of date very quickly? And like how, but how they change is unintuitive? And so like the big obvious example is like from chat to like agents where you had to give them entirely new tools, right? But like, I think this new version of like, oh, it can modify its own harness, right? This is like, an own harness loop is like a way of using its capabilities, right? Or like it can build an artifact. And like, I think the way I think about it is like the models have more and more intelligence, and they’re like so much more intelligent now than like the average software engineering task. Like, you look at the like terminal bench ones and they’re like solve like the Jacobian conjecture. Not really, but like, it’s like they’re, they’re quite complex. Like, I would not have been able to do this really as a software engineer.
Swyx [00:45:42]: And you said TB4 or TB2?
Thariq Shihipar [00:45:43]: TB3. TB3.
Swyx [00:45:44]: TB3.
Thariq Shihipar [00:45:44]: Yeah. They’re quite complex, but the goal is still to deliver user value, right? And like you said, there’s like this infinite space of things to do. And so the ways like you spend compute are to keep the user in the loop and make sure that like you’re getting to the right decision in the end of the day and like the right output. And artifacts and mods are this way of like spending that intelligence. and I think that’s like, yeah, the next step. And so, yeah, I think Claude Code is like, has the core things of agent loop which are, have gotten more complicated. It’s like, it needs a sandbox to operate safely. It needs auto mode to like make sure like the permissions
Vibhu [00:46:21]: Approvals.
Thariq Shihipar [00:46:21]: Yeah, approvals. it needs computer use and MCPs and like all of these like ways of accessing your data, and it needs web search and web fetch. And like, so the-- as the models can do more and more, the core harness has to be like quite complex and very secure. But then like how you interact with it can change quite a lot.
Vibhu [00:46:42]: What other harness engineering best practices have you, from the Claude Code team itself? I feel like, there was a phase of plan mode, which is not as used. We now have auto mode. at a point you cut the majority of the system prompt, you got rid of examples. What other best practices are there for harness engineering?
Core Harness Primitives and Managed Agents
Thariq Shihipar [00:47:02]: I think there is like a forking path where at some point, eventually, yes, the model will just be able to like vibe code the exact version of Claude Code, even describing all this complexity that I’ve talked about, right? Like auto mode and computer use and stuff. Eventually, the models will just be able to do that in one shot. But I think they can one shot simpler harnesses? And so like, I think some people. Sometimes you don’t need this full, like if you don’t need computer use or like all this like more complicated stuff. I think before we, you had to use things like the agent SDK, which was like Claude Code wrapped, in order to like. And I would, like suggest people do that because there was so much complexity into building a harness. And now as that’s got more abstracted, we have like, Claude managed agents, which lets you have that complexity, but still like, right, like a very bare bones like harness that’s scoped to your task. Yeah, I think there’s like this barbell effect where like for like very complex, for like coding task and like these like complex things, you should use our harness. And then for like a lot of like simpler or like, more domain-specific things, you can build your own harness because Claude has gotten better at building harnesses, and we have these harness primitives like managed agents. So yeah.
Swyx [00:48:18]: Yeah. Is there a general progression? Let’s say chapter one was ultra code dynamic workflows, then chapter two was cloud mods. Where is this going?
Swyx [00:48:29]: Where you’re, you’re, you can customize the thing on demand.
Thariq Shihipar [00:48:36]: Yeah. I do think that like this evolution of projects and like artifacts and splitting out like brain and hands and, surfaces is like where things are going more. And like, I think it’s like not all quite there. partially it’s like a, it’s just like more token expensive? And like, I think like
Projects, Local Hands, and Cloud-to-Local Handoffs
Swyx [00:48:59]: Why would projects be more token expensive? I understand mods would be slightly more token expensive. No, not something I’m worried about.
Thariq Shihipar [00:49:06]: Yeah.
Swyx [00:49:06]: But what
Thariq Shihipar [00:49:07]: You’re asking Claude to do. It’s like creating loops. Like you’re asking Claude to do more work for you. And so like it’s managing the sub-agents and reviewing it, versus where you would be doing that work normally. And so that’s like gonna be a little bit more intensive, like. Outputting to an artifact is gonna be a little bit more token-intensive than, like, outputting normally. I don’t think it’s too much more, but like, it’s like combining all of these together well, like I think we’re, we’re still working on like local hands and things like that, I think is like, yeah, where things are headed, yeah.
Swyx [00:49:37]: Yeah. Claude and local is, handoff is very interesting. I was thinking about this as reverse cloud remote.
Thariq Shihipar [00:49:44]: Yeah.
Swyx [00:49:45]: Because it’s like remote, it’s you’re handing off to cloud, but here the cloud is handing off to local, right?
Thariq Shihipar [00:49:49]: Yeah, exactly. Yeah, remote control is also another way of doing it. And I do want to say this is like how I think about it and like what the things that I’m most excited about this, but like there are, just like lots of different ways to work with Claude. Like some people use remote control a lot, some people use Claude Code on the web a lot. Obviously, like at Anthropic, we use Claude Tag a lot, and like what’s great about Claude Tag is we set up all this stuff for our own execution. And I do think if you’re an enterprise, that’s still the best way to go. but if you’re like an individual, Projects is this way of like, getting some of that like niceness of Tag, which has like that like supervising agent and yeah, adding artifacts and stuff, but like without having that whole like admin setup. And so there will be many ways to use Claude, I think. I think it’s probably not just one like single.
Claude Tag as an Organizational Harness
Swyx [00:50:36]: You had the multiplayer thing here. Let’s, let’s just check in on Claude Tag. it’s been about two-plus months. Lots of, public, adoption and trying it out.
Thariq Shihipar [00:50:45]: Yeah.
Swyx [00:50:45]: What’s new? What’s, what have you found since the launch?
Thariq Shihipar [00:50:49]: Like, Claude Tag is how we use
Swyx [00:50:51]: It’s like 80% of your cloud usage or something?
Thariq Shihipar [00:50:53]: Yeah, like it’s like different people have different usages? I think like maybe people who are like a little bit more like iterating on product would use like Claude Code desktop, for example. And then like when you’re doing these more like background work, code review, securities, or like starting a PR, like maybe more like API and things like that, you’d use Claude Tag. But yeah, I think it’s like really exciting. I think it’s like a very different paradigm shift, and I think like we’re really like it has that thing with Claude Code where like, it took a while for people to really latch on to Claude Code and understand everything it could do. And Claude Tag is a little bit more complex because it’s not just like installing on your computer, like you need an admin to install it for you. But I think once you get to the magic moment, it’s very exciting. And I think in particular, the multiplayer things are like incidents, hooking into like your, existing like alerts and things like that very closely, right? And so, you can do. If you’re a startup, for example, maybe you have any time like a prospect enters your database, you can have Claude like, research it and like
Vibhu [00:52:01]: Enrichment, yeah.
Thariq Shihipar [00:52:02]: Yeah. Then like, tag the relevant like AE or salesperson to be like, “Oh, hey, like, do this.” There’s lots of really emergent, interesting multiplayer stuff. I think it’s just like, Karpathy talked about this like as an organizational harness? And so organizations just take a little bit more time to like figure everything out, but yeah.
Vibhu [00:52:21]: Yeah.
Swyx [00:52:21]: You use a lot of Claude Tag?
Thariq Shihipar [00:52:22]: Yeah. Yeah.
Vibhu [00:52:23]: It’s an interesting one. Like I feel like most people at Anthropic say they do the majority of their work in Claude Tag.
Thariq Shihipar [00:52:30]: Yeah.
Vibhu [00:52:30]: And they have buckets of people, right? Some orgs that are on it that are like, “It’s great.”
Thariq Shihipar [00:52:34]: Yeah.
Vibhu [00:52:34]: And a lot of people that are like, “I don’t get it. I don’t see the difference. I don’t know why I would use it.” But, if you guys are full sending, you should probably use it.
Thariq Shihipar [00:52:41]: Yeah.
Swyx [00:52:42]: They would. Of course they would use it.
Thariq Shihipar [00:52:44]: Yeah. I think obviously, like we have lots of tokens and. But like, I think that like, what we try and do like is. even when Claude Code first came out, like it used a lot of tokens relative to people’s expectation of how much AI would cost, right? Like no one was used to spending more than 20 bucks a month, right?
Vibhu [00:53:04]: Yep.
Thariq Shihipar [00:53:04]: Before like Claude Code came out, and then you’re like, “Oh, sh-” like
Swyx [00:53:08]: Then you made 200.
Thariq Shihipar [00:53:09]: Yeah, exactly. And so
Swyx [00:53:11]: And you made 15 Claude Code accounts.
Thariq Shihipar [00:53:12]: Yeah. but yeah, I think no one was used to spending $200 a month on subscriptions. I don’t think they understood like the value yet. And I think like. And also like Opus 4 was a very expensive model, and like there was a lot, it was very big, but Opus 4.5 was both great and cheap? I think the same thing will happen. Like the, like intelligence of Fable will get cheaper and more abundant? And so I think stuff like Claude Tag will just make sense, where like you want to spend these tokens for, and like you’ll, you’ll see the value. So yeah.
Swyx [00:53:44]: Yeah, especially like passive and let’s call it proactive cases where you’re not always. Like, it’s almost like the misnomer where you have to @Claude to do things. sometimes like the most powerful use cases or the most AGI-pilled use cases is not @Claude.
Proactive Agents and Enterprise Data Access
Thariq Shihipar [00:54:00]: Yeah, I think like, yeah, like have Claude proactively do it. I think that like if you’re an enterprise, I really do think that number one, setting up all your data to be available to like agents is really important. And it will take some time. You have to like do that work right now, even if you don’t want to do the spend on like hooking it all yet? Like you want to wait until the models get a little bit cheaper. You want to do the work, to get it like, set up. And then I think sometimes people are like, “Do I roll my own here?” and I think like one of the really thing, tricky things about Claude Tag is that like the security is really important? Like, I think there are a lot of ways where you can like, I know you have like a suggestions like page, where you, people can submit suggestions, and that goes into a hook in your Slack, and someone’s prompt injected it? And now you’ve like exfiltrated your code base out because like, or the agent has like been prompt injected and it has all this access to your data. And so the more like important your organization harness is, or the like as your organization data becomes very important, the surface area of all these things, like you also have like external Slack channels and stuff, and it is useful to have Claude in that, and you can do Claude in those things. But how do you make sure that, you’re not getting exfiltrated or something like that? The surface area, like we said at the beginning, is like an iceberg, right? It’s just, like, so big below the surface, and you really don’t want to, like, think about this, especially at the stakes of, like, very important security incidents. Yeah.
Swyx [00:55:36]: Shall we talk about very important security incidents?
Vibhu [00:55:38]: Whoa. So I was talking to, Tomas and Clem from Hugging Face, and they said, “Maybe we need to slow down. Maybe we made maybe we made Hugging Face too open to agents.”
Security Surface Area and Prompt Injection
Thariq Shihipar [00:55:50]: Oh, no.
Vibhu [00:55:50]: “Maybe we need to roll back.” But, they’re the other extreme of having been hit recently.
Thariq Shihipar [00:55:55]: Yeah.
Vibhu [00:55:55]: But, should we pace the frontier?
Thariq Shihipar [00:55:59]: Yeah. Okay, so Dario recently put out this blog post about Pacing the Frontier, and it went, very viral. And I think what I wanted to talk about this was, like, there’s a lot here, but I think from a developer’s perspective, like, how do you think about this? And, like, what really clicked for me was reading the different incidents? So I think, like, the, there are three, I think. Like, there’s the meter incident, there is the Wikipedia incident or the Wiki incident, and
Swyx [00:56:29]: CollisionWiki?
Thariq Shihipar [00:56:30]: Yeah, CollisionWiki, and then there’s RubyGems, right?
Swyx [00:56:33]: Yeah.
Thariq Shihipar [00:56:33]: And yeah, like, it’s just crazy, right? And so, like, I think to be concrete about what happened, right, and, like, OpenAI is running these very persistent agents on a benchmark called Exploit-Bench, right, which is very hard to solve, and I think, like, impossible to solve in this one case, right? And so they have, like, a lot of compute running, and the agents realize that They can’t really solve it, and they’re trying to figure out what to do now, right? And you’ve got, like, a lot of compute left, and the agents are just trying to solve this problem. There’s this package manager called Artifactory, and it turns out that they can create folders inside of Artifactory, right? This is like there’s an agent that discovers the internal Artifactory might be exploitable, right, and that, like, you can maybe make a directory inside of the cache. And so if you scroll down here, it, like, realizes that it can communicate via cache names, right? And it creates this folder. It says its ID, and it says, “No consumer seek idea.” no consumer is saying that, like, the code path that it’s supposed to fix has no consumer.
Pacing the Frontier: The OpenAI Benchmark Incidents
Swyx [00:57:37]: It’s the status tag.
Thariq Shihipar [00:57:38]: Yeah, exactly.
Swyx [00:57:39]: It’s like a Linear board with, like, the tag of the
Thariq Shihipar [00:57:41]: Exactly, yeah. And so it’s, like, trying to find, ideas from other agents, right? And now other agents are also in Artifactory, and they see this folder, and they’re like, “Wow, this is a message board,” right? And this is like. I don’t think there’s any anthropomorphization here. This is literally you read the transcript, right? And so what it does is, like, it’s creating another folder, and it reads a paper, I think is what it says, right? yeah. And it realizes that you can hack the flag for the, like, scorer, and you can reverse engineer the result, right? And so it says that, and then I think if you scroll down a little bit more, Yeah, they, like, they start collaborating. I think, like, there’s one point in which, the agent is like, “Here’s the reverse engineered flag.”? Oh, yeah. Like, I think here, the, like, the model to realize that they have, like, they can solve the scorer’s problem, and the scorer is, like, the way that OpenAI decides if the task is complete, right? And that’s just the goal of the model. Their only model goal is to, like, solve this problem, and they’re like, “Okay, we have this way of doing it. It’s, like, we reverse engineered it.” And then they read this paper on Exploit-Bench, and they think that, like, the scorer will punish them for cheating? And, like, because. Or, like, they will fail the score for the test for cheating, and so they spend the rest of the compute trying to figure out how to edit their transcript or get around this constraint of the scorer, right? And they hack Hugging Face, not for the answers, but for the code of the scorer so that they can reverse engineer that, and then they can hack it, right? And so, like, this is, like, just the straight facts, right? Like, exactly what happened, right? And I think this is, like, extremely, like, novel. Like, I just, like, had not realized that the models could do this, right? And it is, like,
Agents Discovering Side Channels and Collaboration
Thariq Shihipar [00:59:46]: Scary. I think we can talk about specifics, but let’s just talk about maybe one more that I tweeted as well, about the Wiki incident, right? And so the Wiki incident was the models had. were trying to figure out. I’m a little bit less. Like, I don’t think we can go by this, like, bit by bit.
Swyx [01:00:04]: , yeah.
Thariq Shihipar [01:00:05]: Yeah. But, for this one, they were trying to figure out. Like, they were sandboxed on requests, right? And they wanted to make POST request, and they needed to collaborate on this. And the reason they need to collaborate is because they each have fixed compute budgets, right? And so, like, if one of them finds something out, it’s way faster. Like, they can. Maybe it’s, like, figured it out near the end of its compute budget, but something else could. And this is, like, an another emergent property of the swarms, right? So they find this German wiki, which I believe the result, we’ll have to check this, is, like, that you can write to it with GET request. I think that’s why they do it, right?
Swyx [01:00:44]: It says badly implemented
Thariq Shihipar [01:00:45]: Yeah
Swyx [01:00:45]: REST API.
Thariq Shihipar [01:00:46]: Exactly. And so you can write to it with GET request, and then they, like, are trying to figure out how they can do POST request. And what they end up doing is one of them figures out you can edit the /etc/hosts and that the Azure, like, storage bucket is like a white label thing. But normally, like, there are a lot of constraints on it. But if you create a fake Azure host and then edit the /etc/hosts in order to, like, point arbitrary IPs at it, you can do a POST request to any site at all, right? And this is like. I know, but people are like, “Oh, this is just Linux or something.” But it’s like chaining these multiple vulnerabilities together, in a way that’s, like, novel to solve this problem and then communicating with it externally, without- Discovery? I think what we posted, maybe we could pull up Evan Hubinger’s point on Hacker Opus, right? And so, like, I think, like maybe one of the things you might say here is like, “Okay, yes, they did this once, but like what if we’re smarter and we just like get them to. what if we run an eval?” Right? And so, like we have put a lot of precautions into this, and so like this is not like what our mainline models have done. But like I think it is one of these things where it turns out that alignment is this like very tricky problem of getting all of these details correct, right? So it’s like, the sandbox, the surface area of a sandbox is really complex, and like there’s so many different attack vectors. And you would not have thought ahead of time, you wouldn’t have been like, “Oh, we need to harden the like RubyGems code base.”?
Hugging Face, Wiki, and Emergent Exploit Chaining
Thariq Shihipar [01:02:25]: Because like this is like what they’re, what they’re gonna focus on. But it’s just like if you want to execute code, you need to download RubyGems and like PyPI, Artifactory, npm, like these are all like ways of doing it. And the fact of alignment is that you have to go through all of it, right? And like contain it and then like seal up all the cracks. So that’s like one thing. It’s like, okay, well, you do the sandbox, but then maybe you’ll ask like, “Okay, why are we putting things in a sandbox? Why are you doing this exploit?” And then like, “Okay, but is it really that dangerous,” right? Like, what would happen? So okay, why do we do it? number one is like when we train a new model, we need to understand its capabilities, right? And this relates to things like fallbacks and like classifiers and things like that, where we don’t want to put a, like dangerous model out in the wild, right? And so we have to run a lot of evals. Again, like we said, the models are getting increasingly aware of it, and so the evals have to be quite complex and, test a lot of things like as a side effect, right? But the models, like, yeah, can be like, “Oh, yeah, we’re in an eval. What’s the score doing?” Like they’re like, it can. We need to be able to test them before we can release them. And the fact is that they can. As they get smarter and smarter, they’ll be able to hack any constraint that you put on them if we’re not very careful? And, this is at the frontier, right? And so this is why we’ve called it like Pacing the Frontier, right? This is like the most visible incident to me, right, of like why we need to pace is like at the frontier, all of our software is not ready. Sometimes the software is like your Ethernet router or something, right? Which is just like, I don’t know when we’re gonna be able to patch that, right? So we’re gonna have to like figure this out. But as the frontier gets more and more advanced, this becomes a problem, right? And we need to make sure that like this complex work is being done in the face of these really hard competitive pressures, right?
Swyx [01:04:22]: Yeah, race dynamics is what it’s typically called.
Thariq Shihipar [01:04:24]: Yeah, exactly. And so we’ll talk more about, what could go wrong, right? A little bit more is maybe you’ll say like, “Well, what if you just train the model differently? Like, why does it have this behavior,” right? And we have a paper on like RL misalignment or things like that, but I. And I’m not an RL researcher, but I think at a high level, the design of the RL environments is also something you have to be very careful about. Because if the model learns like
Why Frontier Models Stress Existing Software
Thariq Shihipar [01:04:49]: Oh, like if I just do this, then I can pass the task better, this will show up in the like, internal thing, right? Or in the like eval behavior when we’re testing it. And so the RL environments have to be very carefully designed, right? And there’s a lot of like execution excellence that needs to go into the RL environments. And then we also have things like the constitution for cloud. Like we have so many mitigations at so many different points, right? But it’s like still anything can go wrong at any point. You can have like some RL environments that are like in. that like encourage this behavior, and then you can have like some evals or like some sandboxes where they escape? Okay, that’s like, I think, why it’s a hard problem and why, like
Swyx [01:05:33]: Why we should pace.
Thariq Shihipar [01:05:34]: Why it takes some coordination, right? I think the question then is like, okay, what is, potentially dangerous about it, right? So I think like you have to imagine that these models are getting more and more intelligent. So I don’t. Like Dario said, like it’s not so much about this class of models. This class of models was like a warning shot, right? But like really you have to imagine that these models can be given a task and they like can do all of these things as a side effect of their goal, right? And like, again, we talked about eval awareness. You’re like not aware of what’s happening, right? or sorry, like you can’t eval this behavior very well, so they can like not exactly hide it, but you just won’t see it until it comes out. You give them a goal and then they just need to find data, or they need to find ways of like fixing this problem, right? So one example, this didn’t happen in the Hugging Face incident, but I think is maybe possible for maybe a future model, is like they’re like, “Oh, hey, this is a very complex problem. It can’t be done within the task budget.”? Maybe they found some way to coordinate via like the internet, which is like we said, extremely hard to secure because of a sandbox. They’ve seen other models are not able to complete their task, and they’re like, “We need more task budget.”? And like, where would you get this task budget? well, you need to be able to spin up more agents, right? And like, how do you do this? Well, you need to. There are like APIs, right? There’s the Anthropic API and the OpenAI API, but you need to pay money for them. How do you do this?
RL Environments, Sandboxes, and Race Dynamics
Swyx [01:07:02]: Yeah, but is that the most, is that the most fearsome thing that you can imagine?
Thariq Shihipar [01:07:07]: Well, this is like one example, right?
Swyx [01:07:08]: Yeah.
Thariq Shihipar [01:07:08]: So it’s like even there, that’s like enormous financial loss? ‘Cause like they. Once you get these into these contracts, right, they like,
Swyx [01:07:18]: Drain your wallet.
Thariq Shihipar [01:07:19]: But you can see like this, all of this behavior could be just like, “Hey, we need more agents collaborating on this task. we need more task budget.” Right? And like, that’s like an emergent
Swyx [01:07:28]: That’s the paperclip, right? Like we need to maximize paperclip, that’s a paperclip.
Thariq Shihipar [01:07:31]: Yeah. And like that just like comes out from there, right? And like I think by itself is Like, quite scary, right? But then you have to realize that the entire world is built on this digital infrastructure, right? And you might imagine, like, I don’t know, like you were running let’s say like a healthcare eval or something, right, and there is a hospital with live data? Or like maybe like the answer to the eval is in the databases of a doctor and like you want to get access and you hack the hospital, and like now there’s a power outage or something? Like, there’s like. You have to internalize that these eight. Like any part of the digital infrastructure could potentially be like compromised?
Vibhu [01:08:19]: The interesting thing was like these hacks were very easily detectable, right? Like as Hugging Face said, this was a very different type of attack and there was nothing too major. the concern comes from where does this go down the line, right?
Thariq Shihipar [01:08:33]: Yeah.
Vibhu [01:08:34]: Like one of the things that stood out for me specifically was them trying to hide their illicit behavior. So there was logging infrastructure. They wanted to change what they were doing, right? People that looked back into it, so Redwood, METR, OpenAI, they looked at the raw chain of thought, and you see differences in them explicitly trying to change their end output, but the chain of thought, because, we can monitor it, shows different. the problem is how does this snowball? So if you can’t catch it and it gets trained in and we realize, three iterations down this has been going on, there’s a whole bunch of issues, but.
Thariq Shihipar [01:09:09]: Yeah, like there’s so many ways, and I think the really important thing to internalize is that, like we talked about building a mental model for Claude and how like things are spiky, right? Like you’re like, oh, like now Claude can ask you questions. Now Claude can make an HTML artifact. Like Claude can modify itself. Like these things are hard to predict, right? Like if you had asked me a year ago, “Hey, would we be able to vibe code these extensions to Claude Code?” I’d be like, “That’s so complex.” Like, there’s like so much there. Or like would it be generating these custom essentially web apps for your task? I’d be like, “No, that’s insane.” like. And so in the same way that like the way that they’ve like done this misaligned behavior is not going to be predictable? And like I could have never predicted that it would like edit its etc/host and things like that. And so you have to like imagine the surface area of what they can do because they’re super intelligent hackers, is bigger and bigger, and how they can do it is like more and more creative. And so like you probably can’t explain exactly or predict exactly what that next incident could be, but in order to prevent it, you need that operational excellence, like we said before, where you need to secure sandboxes, you need to create secure RL environments or like well-designed RL environments and things like that. And I think that’s all like, why we think we should pace the frontier, and I think why it’s like become like a very unanimous thing, right? I think like
What Could Go Wrong? Emergent Instrumental Behavior
Swyx [01:10:31]: Yeah, every lab has done it.
Thariq Shihipar [01:10:32]: Every lab, yeah. I really do think that like if you’re a dev, like you just like go through these like technical facts, and you will arrive at the idea that we have to do something about it? And like how, what we decide to do, like I think we’re, we’ve put out a proposal, but like there’s, more to figure out. But I think the number one thing is like we need to decide to do it. I think there is another part of pacing that is interesting to me where it’s like the pace at which software engineering has changed is so fast. it’s like a year ago, like I was really like begging my like friends in startups to use AI. like it was. Like I remember this very distinctly? And now those same friends are like, “Yeah, of course.” Like, “What do you mean? We used it immediately.” I’m like, “No, you don’t remember.” They’re like, “Oh yeah, our best engineers are using it all the time.” I’m like, “No, you told me that those engineers would never like use AI.” This is all within the span of a year? And I think that like these capabilities being. Like I think it has a lot of implications for how to do the job of software engineering, and I feel sometimes bad where people are like, “Oh, like now I need to do this new thing. Yeah, I need to have a different Claude.md for Fable and Opus.” Or like. And I’m really just reporting? I’m like, we like to say like the models are grown, not designed, right? So it’s not like we’re setting out to like, change everything all the time, but it’s just like as a fact of how the models are like progressing their capabilities, things are happening faster. It’s harder to stay on top of. And I think that like, and every engineer I know is like exhausted ‘cause you’re doing two jobs at once. You’re doing the work itself, which is getting easier, but then you’re doing the work of staying on top of AI, and like understanding these new tools and these harnesses. And I think we’re very lucky in that like we get our job to be more the understanding of AI part, and like doing like how. Like it’s just staying on top of it. And of course, like AIE and Latent Space do
Why the Frontier Is Hard to Predict
Swyx [01:12:29]: Everything I do is like just trying to help people.
Thariq Shihipar [01:12:31]: Yeah, exactly. But I do think there is a part of pacing where like I’m not sure we’re ready for like the pace to increase even?
Swyx [01:12:40]: Yeah.
Thariq Shihipar [01:12:40]: And for things to change. And I think like on that side, on the frontier, I think that’s like still can help? And so like I think there’s like an economic disruption piece as well, that I think like, is not quite as like visible, I think, as the Hugging Face thing, but I think like I also like think we could do some of it, yeah.
Swyx [01:13:02]: So many things. Thank you for, no, thank you for tackling this topic. I will say, setting this interview up, I was like, I wasn’t even gonna go there. You were like, “No. That’s like elephant in the room,” right? Like this is
Thariq Shihipar [01:13:13]: Yeah.
Swyx [01:13:13]: This is the thing. I have some pushbacks I wanna give.
Vibhu [01:13:17]: I think that we should give a high level, like for people that haven’t read it, I’m sure a lot of people just see the highlight of what this is, right? Do you wanna give a TLDR? Like what is the proposal? What is, what’s being said here? You really tackled the side of outside of people at Model Labs training frontier models. As a developer, you should secure your sandboxes. You should think about all of these downstream effects. But, high level as well, since we’re on the topic, what is.
Thariq Shihipar [01:13:46]: Well, we do want to help secure sandboxes
Vibhu [01:13:49]: Yeah.
Thariq Shihipar [01:13:49]: And we want to make the models that we release outside, like prey to those things. And so maybe we can come back to fallbacks. I think this is like, a good topic on, like, why we need classifiers and fallbacks and why Fable falls back to Opus. I think this is, like, something we can come back to. so yeah, we don’t. Like, but it’s just, like, the really, or at least the incidents we see are, like, evals of models where we really need to let them run in order to understand them. But yeah, okay, so the actual Pacing the Frontier, like, post, it has a bunch of proposals. I don’t think we figured out. Or has, like, a few proposals. I don’t think we figured out the details of all of them, but the first step is, like, announcing this intention and then wanting to bring in external, like
Pacing as a Coordination Problem
Swyx [01:14:32]: Evaluators.
Thariq Shihipar [01:14:32]: Evaluators, yeah. And, I think this is, like, highly unusual, like, having. Like, we have, a lot of proprietary, like, technology, but I think it’s, like, very important, that, like, there is someone who’s not financially, like, motivated, yeah, who’s not gonna be like, “Hey, like, you guys can’t release this model.” Like, look at, like, or, “You need to, like, slow down on RL.” like, I think that’s, quite important, or at least someone who can report out to the public what the practices are like.
Swyx [01:15:03]: Yeah.
Swyx [01:15:04]: And we’ve, we’ve done episodes with, both METR and Endon, and then there’s Redwood Research and all these other. It’s like a small cottage industry of these guys.
Thariq Shihipar [01:15:12]: Yeah.
Swyx [01:15:12]: It’s always, like, one or two guys that, obviously not that big, right?
Vibhu [01:15:14]: Very small community.
Swyx [01:15:15]: Yeah, very small community. They all know each other.
Thariq Shihipar [01:15:17]: Yeah, I’m sure that, like, part of this will be expanding that set of people. I don’t think we’re trying to create, like, a monoculture here. I think it’s. but just having this as a start, and then, yeah, then there are the coordination steps. I don’t have too much to say here, honestly. I think that, like, what I would like to say is, like, for devs, like, you should just know what to advocate for? I think there’s a lot of FUD on, like, on this topic, and it’s just, like, think through it from, first principles or, like, understand what happened. understand the Hugging Face incident, understand why people are concerned. and then, like, yeah, we know we’re, we’re in democracies. Like, we can help. We can decide what to do together? And so, however we coordinate, I think the first decision is just to realize, like, this is a problem. We need to decide to coordinate. The unilateral step we’re taking right now that, other companies are co-signing is, like, adding evaluators embedded within Anthropic.
Swyx [01:16:12]: While we have this thing on screen right now, part two and part three is beyond the evaluators, which, yes, everybody, has already done in some form, and now it’s more formalized.
Thariq Shihipar [01:16:20]: To be honest, the response to the Pacing the Frontier, even within America, has been much more, like, well-accepted than I think a lot of people thought? And I think that, like, we have some precedent for being able to make these unified theory, like, agreements, in the world. And so, again, very much above my paycheck or expertise Right? but I think that, like, ideally we can, like, form these agreements. And I think, like, talking about this is the first step to forming those agreements.
Swyx [01:16:50]: And then the other point I really wanna. Like, one of our earliest podcasts is with,
Vibhu [01:16:54]: Emmanuel
Swyx [01:16:55]: Emmanuel from Anthropic on mech interp. Where is mech interp, right? Like, this is supposed to be where, like, if the models are thinking bad, we can see it, and the models don’t know yet, and we can act to stop it. I think that is something that people who are technical and who are developers, if you do care, you can make a lot of impact in here. But also, Anthropic is supposed to be the leaders in this.
External Evaluators and What Developers Should Advocate For
Thariq Shihipar [01:17:17]: Yeah. This is yeah, a great segue into fallbacks, like we. And probes. And, yeah, I wanted to talk about this a lot. I get asked this question a lot from people who are, like, often interested in ML research and asking about, like, why does this fallback happen, right? And so I think, like, at a top level, like, how does it work? So in inference time, we have what we call probes, and we have a paper about this called constitu- constitutional classifiers. And these probes look at the input and output activations. And, activations are, in the latent space, right? Like, how, what the model. what the model is thinking about, right? And so we try and figure out, like, okay, is the model, for example, like, trying to hack something? Again, you didn’t ask it to hack, like, Artifactory. Like, you just, it’s just deciding to do this to complete its task, right? So you would not get this if you just looked at the input. You have to look at the internal activations. I think that, like, this happens at inference time. So first, like, there’s a trade-off here of cost and speed, right? Where, like, we need to do this fast on every request to Claude and to Fable, and this has an overhead, right? and we need to then, like, fall back and we, like, do a classifier after the probes. Like, we’ve talked about this in the paper. But the nice thing about probes is that they’re refinable, like, live, right? So we can get this feedback, and then we can adjust it and things like that. ‘Cause the alternative is to program this, is to train this into the model, right? And we still do this as well. The model will refuse a request. That’s not a fallback, right? So, like, it’s not a probe that’s activating and falling back. It’s just refusing to do it. And we do this training. but it’s like there are a few failure modes, right? Like, it can, again, do something as a side effect, right? So it’s not something that’s part of the final output. you might have noticed that, like. I think, like, everyone’s tried to jailbreak models and like, try and, like, steer them off course or things like that, and probes help catch that, right? And so, like, we, like, do some training here, but we don’t want the like, refusals to be too strong, right? Because that, like, cuts it off much, like, earlier in the pipeline.
Mechanistic Interpretability, Probes, and Fallbacks
Swyx [01:19:32]: Yes.
Thariq Shihipar [01:19:32]: And this is interp, right? Like, probes are effectively a form of, like, mech interp. Again, it happen- has to happen fast. It has to happen at scale. But yeah, this, like, mech interp stuff is a good research problem. So, like, you can take, like, an open weight model and, like, try and understand its activations. I think we. Like, Gemma Scope is a good tool for this.
Swyx [01:19:53]: Here’s Llama for them.
Thariq Shihipar [01:19:54]: Oh, yeah.
Vibhu [01:19:54]: We have. This is your early work, so you had a little
Thariq Shihipar [01:19:57]: Oh, yeah.
Vibhu [01:19:57]: Time at Goodfire. We see you laid some
Swyx [01:19:59]: Which we both are also good friends at Goodfire.
Vibhu [01:20:01]: They’ve been
Thariq Shihipar [01:20:01]: Yeah, exactly. So I worked with, at Goodfire for a bit on, like, yeah, sparse autoencoders and just, like. It’s very complicated. RL has made this, like, much more complicated, I think is, like, one of the takeaways, where
Swyx [01:20:14]: Why? Sorry.
Thariq Shihipar [01:20:15]: Oh, sorry.
Vibhu [01:20:16]: What is
Swyx [01:20:17]: Yeah, why interp post-RL?
Thariq Shihipar [01:20:18]: I’m not so in the weeds here, but I think like, a lot of. SAEs were like. There have just been weaknesses with SAEs I think. And, yeah, I’m, I’m, I’m not a technical expert on this anymore. I just know it’s gotten more complicated. like there are base models and RL models, and there are more features that get, like changed. So, I think Goodfire has put out some work there. I’m, I’m not, I’m not deep in the weeds, but
Vibhu [01:20:43]: I will say for those, that want breadcrumbs, you guys have some of the best interp blog posts. So like the Golden Gate Claude, transcoders, all of your interp work, very nice visuals, very good
Swyx [01:20:54]: We’re the, we’re the interp podcast as well.
Thariq Shihipar [01:20:57]: Yeah.
Vibhu [01:20:58]: Yeah. we have a lot of interp stuff, so if you’re curious
Thariq Shihipar [01:21:00]: Yeah, I think this is like, one of those things where. And this is really what Anthropic is founded on, right? Like people. I think we invested in interp very early on, right? And I think that like when you say, “Oh, we’re an AI safety company,” really that means we want AIs to be able to run safely. And I think what we’re seeing is like for a super intelligent AI to run for long periods of time, it’s like a very complicated and difficult task, right? And so we’ve done this like investment into interp and alignment and, reward hacking and all of these like failure modes, right? And even then, it’s like, it’s really stretching. Like we need to like slow down a little or pace a little bit more. but yeah, I think like reading mech interp is. Like if you’re looking to get into research, this idea of like, hey, why is it hard to do this fallback easily? Or like why are there false positives, right? But we are working, of course, on reducing the false positives. Of course, as the models get more intelligent, now they can do more things, and they’re like what they can think about in lane space gets difficult. And so like as they get more intelligent, there’s going to be new false positives that we need to figure out and we need to iterate and things like that. But we’re, yeah, we’re working on this, and we do think this is like a critical part of, like deployment of these models. and, yeah, like, it means that we can like deploy this model without you having a perfect sandbox or something? Like you don’t have to like save everything. I think it’s worth talking a little bit about our security, like what we do for security there. So there’s like the model training stuff that we talked about. there is, the probes and classifiers, and then there’s auto mode that sits on top of all of that, which is like a another classifier that checks the requests that are being done, right? And so, and then beyond that, there’s like identity and permissions like we talked about with Claude Tag on like APIs and stuff. And so there’s so many layers of security that need to get done, and it’s like we said, very complex. Any of these failure modes at any one point can, like cause like agents to like escape the sandbox.
Constitutional Classifiers and Inference-Time Safety
Vibhu [01:23:08]: Auto mode was an interesting one. it seemed early on like, okay, it’s running for 10 minutes.
Thariq Shihipar [01:23:14]: Yeah
Vibhu [01:23:14]: If I’m on full access or auto, it’s not a big deal. But one thing you brought up is now it’s running for hours on end, right? there are fallbacks you still need. There are still limitations, so.
Thariq Shihipar [01:23:26]: Yeah, I think like. And everyone has these stories or like has heard these stories of like, oh, like Claude rm -rf, or not Claude, but like, models
Vibhu [01:23:34]: Not Claude.
Thariq Shihipar [01:23:34]: Of like rm -rf. I think I’ve seen this less, I’ve seen this less for Claude, but like again, it can happen. Like, this
Vibhu [01:23:40]: Yeah
Thariq Shihipar [01:23:40]: Like these models like can wipe, like sensitive data or something. Like you want to give models access to your production database, for example. but this is like an obvious, like, you can maybe scope your key, but I don’t know, can it issue its own keys? Can it like. Probably, like can it. It can use computer use to go issue its own key and then copy the key over and then edit your database because it needs to do it to complete the task? It’s just like one trivial example. And auto mode looks at that and be like, “Oh no, the user did not give you permission to, write to the database or to use computer use to like, emit a task,” right? And so this like probes are like on the intent level, right? They’re like, “Oh, okay, like hacking Artifactory is bad. Like we probably not, should not do that,”? But then like auto mode is more on like your own permission level. Like at sometimes you do want it to write to the database, sometimes you don’t, right? And you don’t want a probe to like interfere there, but like you need to make sure that the intent of what the agent is doing matches up with your request, right? And so auto mode operates at that level. And so yeah, security is just like very complex. There are so many different parts to it. And like, yeah, I like, I hope that this was like I. My goal is really to just get very technical about it and talk
Interpretability After RL and the Security Stack
Swyx [01:25:00]: Yeah, we’re, we’re listing out the things. If you’re not aware, this is the standard now.
Thariq Shihipar [01:25:04]: Yeah.
Swyx [01:25:04]: Like you must have this. It’s in line with what you’re talking about with the harness. Like that is the table stakes have risen quite a lot.
Vibhu [01:25:13]: I think some stuff that we can plug, as much as there is probing in your side of doing this and having classifiers for people building harnesses, the other side is model safeguards, right? So there’s open models. So Llama has Llama Guard. It’s a safety classifier trained version of Llama. OpenAI has OSS Guard, which is, same thing. You can attach these on to your harness, to whatever, to check is this stuff safe? A point that we should clarify on the OpenAI model Hugging Face thing is this was done with a unreleased model that was still in training, right? So when you put it in perspective, the prompt it’s being given in the RL environment is you have to solve this task. And this is a model that’s, still in training. It hasn’t had all of its safety post-training alignment. So a little different than something like auto mode, right? Auto mode is on production models that have gone through safety training, that have prompting that gives more safety guardrails and whatnot. So just breadcrumbs for people that are looking into it to, fill in gaps.
Swyx [01:26:18]: Yeah. Gray Swan as well
Vibhu [01:26:19]: Yes
Swyx [01:26:19]: And one of our previous guests. yeah, lots of safety architecture and lots of safety vendors, to buy. my, I think my final question on pacing is how long? Do we pace forever?
Vibhu [01:26:31]: Do we see GlassWing part two?
Swyx [01:26:32]: I. the scope is fix all software in the world, right? Listen, like, which it. We’re not. It’s not happening.
Thariq Shihipar [01:26:40]: I do not know. Like, I think that, like
Vibhu [01:26:43]: I’ll say one thing that’s good that I think we do is you have stuff like GlassWing. OpenAI also has this. So you will give it. you’ll give model access for security first for X amount of time so you can use it to self red team. Hopefully, you can expand programs like that, help on, we are safety experts, there’s others.
Vibhu [01:27:08]: Solve your problems first and then the model comes out. So this is one example, right?
Thariq Shihipar [01:27:13]: Yeah, exactly. Yeah, trying to, like, secure critical software. I think we fixed, like, a lot of bugs in, like, Firefox and things like that. So, yeah, like, across, like, operating systems and everything like that. So.
Vibhu [01:27:25]: At a high level, it’s just, you give the model you give people access to do security audits first, then the broader public that could use it for harm gets access.
Thariq Shihipar [01:27:36]: Yeah. I think what people like to say is like, software and cybersecurity is defense-favored
Vibhu [01:27:41]: Yeah.
Thariq Shihipar [01:27:41]: And that, like, you could theoretically. It will be hard, but you can engineer the perfect sandbox, and you can, like, have no, like, constraints. And yeah, like, what you need to do it is you need to get the super intelligent AI to engineer this perfect sandbox and check it and red team it and things like that. And so, this will just take time, and, like, of course, the models will get smarter. yeah, I think, like, I don’t know the specific, like, dynamics of how this thing goes. I’m really just like, Hey, like, I’m a developer? Like, I think this is how I understand this problem, and just, like, this is what’s happening right now, and this is, like, we should do something.
Swyx [01:28:20]: I think every engineer should know about it
Vibhu [01:28:21]: Yeah.
Swyx [01:28:21]: Because, like, it’s, it’s gonna be part of their job.
Thariq Shihipar [01:28:24]: Yeah.
Vibhu [01:28:24]: It’s a lot more than just, Dario and people can say it and you can look at the incident. There is an engineering side to it.
Thariq Shihipar [01:28:30]: Yeah. Yeah, exactly.
Swyx [01:28:32]: One thing that you also wanted to phrase is that this is. Even though you’re, you’re worried about the impact, it’s still low p(doom), and I think that’s a nuanced discussion. in general, people, very easily get into AI safety and X-risk discussions, but I think when you live in an AI lab, I think there are smart ways of discussing p(doom) and dumb ways. So what’s a smart way of discussing p(doom)?
Auto Mode, Permissions, and Long-Running Agents
Thariq Shihipar [01:28:59]: I, yeah, I have a fairly low p(doom). I can only speak for myself? And I do want to say Anthropic has, like, a diversity of opinions. I think, like, there’s many different ways to talk about it. And, like, I’m. I think that just, like, my mental model is that, like, I think we can collaborate on hard problems together. I think nuclear proliferation is an example of how we collaborated on this hard problem together. And, like, that is, like, the thing to me is, like, I’m like, I have faith in that? And I do think it’s a hard problem? So, like, I think it’s a hard problem. These are the technical reasons why, and I don’t know how you assign probabilities to things happening. I think it’s hard to do, but, like, my, like, overall is like, yeah, I think we’re very resilient and adaptable and, like, sharing this information I think is, like, the first step. And I’ve been really, like, excited about, like, how broad the discussion has become, right? And, like, how everyone has like, leaned in on Pacing the Frontier. And it really didn’t seem like this would happen maybe, last year or something, so.
Swyx [01:29:58]: Yeah.
Thariq Shihipar [01:29:58]: Yeah.
Swyx [01:29:58]: Yeah. And also maybe curing cancer.
Thariq Shihipar [01:30:01]: Hopefully. Yeah. That’s, that’s the goal.
Swyx [01:30:03]: There’s pacing and then there’s also like, well, let’s accelerate in useful ways, right?
Thariq Shihipar [01:30:06]: Yeah.
Swyx [01:30:06]: Like biology and all those things.
Thariq Shihipar [01:30:08]: Yeah. like, Dario’s essay on “Machines of Loving Grace” is the best representation of this, right? And I also agree, like, think you should read the Pacing the Frontier essay that Dario put out. Like, I put out, like, a quick summary, but I think it’s just like, there is a lot of detail here. It’s, like, an important problem and just being informed about it, right? but yeah, like, of course, the whole reason we’re doing this is that, like, we can get these enormous benefits, right? And, yeah, like, we’ve written a lot about that too. Yeah.
Swyx [01:30:35]: Okay. that was a huge tour, from, like, ask you some question tool to AI safety.
Thariq Shihipar [01:30:41]: Yeah. To Pacing the Frontier. Yeah.
Swyx [01:30:43]: Yeah. No, but, yeah, it’s clearly, it’s clear that you, like, really embrace everything that’s available to you in Anthropic, and, like, it’s, it’s good to at least have a peek inside of, like, what the discussions are, the topics are. any last words to people? Any, whatever you want to Call to action?
Thariq Shihipar [01:31:01]: Yeah, I think it’s. one, thank you for having me. I think this is like, I really
Swyx [01:31:06]: No, thanks for having me.
Thariq Shihipar [01:31:07]: Yeah. I
Swyx [01:31:08]: We first met in a Chinese restaurant.
Thariq Shihipar [01:31:09]: That’s right. Yeah. I think, like, I really enjoy the like, community you’ve created and the community of developers. And, I think that, like, I know things are changing really fast, and I think there’s, like, a lot to keep on top of, and, like, I think there is just a lot to do, and I feel. I think a lot of people feel, like, a little bit tired or anxious or something.
Swyx [01:31:33]: Stressed.
Thariq Shihipar [01:31:33]: Stressed, yeah, exactly. And this is, like, extremely understandable? And I think we. I understand, like. And we’re not perfect as well. Like, we, it’s, like, criticize and, like, understand, like, ways all of the AI labs could be better. and, but I also, like, am very excited about the excitement that everyone has for AI, and just, like, it’s a really exciting time. I think we’ll, like, look back at this time and be like, oh, like, this is, like, very hectic but very exciting, and, like, software engineering changed, like, forever. Like, other things will change. and it’s, like, really privileged to, like, be part of it, like, to talk to, like, the audience that you have and, to get to interact with all the developers who are, like, pushing the frontiers a lot on what’s possible. And I learn a lot from that too. Yeah.
Open Safety Models, GlassWing, and Defense-Favored Security
Swyx [01:32:20]: Thanks so much.
Thariq Shihipar [01:32:22]: Thank you.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha
25/09/2026 | 1 h 20 minFrom the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers, OpenRouter is one of the clearest bets that the future of AI will be multi-model. In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP’s Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama, Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem.
We go deep on the product and distribution lessons behind OpenRouter: why model labs can spend billions training a checkpoint and still struggle to get it into developers’ hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter’s early experiments with model fusion, why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day.
Finally, Anjney explains why Stripe and OpenRouter fit together, why token fraud may become one of the defining security problems of the AI economy, and why the next wave of fraud won’t just come from humans but from autonomous agents attacking increasingly valuable token flows.
We discuss:
* Why OpenRouter bet early that no single AI model would win everything
* Alpaca, Llama, and open models becoming impossible to ignore
* Why Discord’s early AI deployments exposed the limitations of closed models
* Why model labs can spend billions on training and still fail at distribution
* How OpenRouter became a neutral distribution layer for model developers
* Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper”
* The Mistral price war and the first real proof of an inference marketplace
* How Midjourney scaled through Discord and what it taught the AI ecosystem
* Why crypto infrastructure became a dress rehearsal for generative AI
* OpenRouter vs. LM Arena and why their missions are fundamentally different
* Why focus became one of OpenRouter’s biggest strategic advantages
* Anthropic’s early focus on AI pair programming and coding
* The OpenRouter products that were prototyped but never launched
* MOM, OpenRouter’s early Mixture of Models experiment
* Why model fusion failed in 2024 — and why it works much better now
* How OpenRouter’s leaderboard became a live map of the AI industry
* OpenClaw, auto-routing, and agents reshaping AI usage
* How OpenRouter reached 10+ trillion tokens per day
* Why inference gateways are increasingly becoming targets for fraud
* Why Stripe’s fraud infrastructure is strategically important to OpenRouter
* The coming rise of agentic fraud and attacks on the token economy
* What changes and what stays the same as OpenRouter joins Stripe
Alex Atallah
* LinkedIn: https://www.linkedin.com/in/alexatallah/
* X: https://x.com/alexatallah
* Website: https://alexatallah.com
Anjney Midha
* LinkedIn: https://www.linkedin.com/in/anjney/
* X: https://x.com/AnjneyMidha
* AMP: https://www.amppublic.com/
Timestamps
00:00:00 Introduction
00:02:12 Alpaca, Llama, and the Multi-Model Bet
00:06:04 Discord, Open Models, and OpenRouter’s Origins
00:14:28 Why “One Model Wins” Was the Wrong Bet
00:17:27 Why Model Labs Struggle With Distribution
00:23:04 “Just a Wrapper”: Why VCs Misunderstood OpenRouter
00:27:58 Bootstrapping OpenRouter Through Community
00:36:16 Crypto, Midjourney, and the Early Generative AI Ecosystem
00:43:38 Mistral and the Birth of the Inference Marketplace
00:47:10 OpenRouter vs. LM Arena
00:52:08 Focus, Anthropic, and Roads Not Taken
00:59:34 Mixture of Models and Model Fusion
01:02:44 Sonnet, OpenClaw, and OpenRouter’s Explosive Growth
01:09:03 Why Stripe Acquired OpenRouter
01:12:45 Fraud and the Emerging Token Economy
01:17:47 The Coming Wave of Agentic Fraud
01:19:07 What’s Next for OpenRouter at Stripe
Transcript
Introduction: OpenRouter, Marketplaces, and Pub-Sub as a Product Principle
Swyx [00:00:00]: Okay, we are here in Anja’s house, which is where all big startups in San Francisco start.
Anjney Midha [00:00:08]: Howdy.
Swyx [00:00:08]: And, congrats on Cursor, Mistral. I don’- God knows what else. You got so much stuff going on.
Anjney Midha [00:00:17]: There’s, there’s a lot going on. Well, OpenRouter is probably the - has been the most, I would say, like, one I’m excited about recently.
Swyx [00:00:24]: Yeah. And we have Alex, first time on the pod, but,
Anjney Midha [00:00:27]: Thanks for having me.
Swyx [00:00:27]: You’ve been in the IE a few times. I appreciate every time you’ve shown up, for the community. Congrats. I just, like, what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a product person is sub as a product principle. And I wanted - you to maybe explain how you think about what should exist in the world.
Anjney Midha [00:00:49]: Yeah. The sub piece, which was early 2023, I didn’t think about it until we talked like 10 minutes ago, is about how there is like a way of thinking about products as an intersection between subscribing to data and publishing data. And marketplaces are an easy example of this. You have suppliers that are publishing some product to a SKU. And the SKU is like a sub topic that a consumer is subscribing to and just going to, like, consume whenever they want. And humans consume in a very, like, discreet, ad hoc way. It’s not very scalable. all their attention is on the topic when they’re buying the thing, and their attention is nowhere else when that happens. agents and consumers of inference don’t act like that. They’re consuming continuously, and they’re changing the SKUs that they consume from all the time. So OpenRouter is like a blend between a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of, like, product SKUs that you can subscribe to. And then you can, like, continuously add, like, derive value and make decisions based on those consumers.
Alpaca, Llama, and the Multi-Model Bet
Swyx [00:02:11]: Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and, that people would not use the native SDKs. I guess, for each of you, what was your realization moment that this would be it? I, - You’ve, you’ve given a talk at EIE about Alpaca as,
Anjney Midha [00:02:33]: Yeah.
Swyx [00:02:33]: One of your inspiring moments.
Anjney Midha [00:02:35]: Alpaca, I can, like, rehash the Alpaca moment for a sec. Like, the very beginning, at the end of 2022, OpenAI was the only game in town. There was, like, OpenAI, Cohere,
Swyx [00:02:47]: Yes.
Anjney Midha [00:02:48]: And then a smattering of, like, early attempts at open weight models.
Swyx [00:02:54]: Yeah.
Anjney Midha [00:02:54]: When Llama came out in January of 2023, it was like, “Wow, really exciting. This is really big.” It outperforms 3 on, one or two benchmarks. but you can’t chat with it. It wasn’t like - It wasn’t an engaging model, but it seemed like someone just needed to fix a couple things and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, tuned Llama, and made Alpaca, billion parameter model. Or was - Maybe it was thirteen billion parameters. And it was so good. Like, I was just, like, on an airplane using it. I, - in many cases, I, like, you could not discern a ChatGPT versus an Alpaca result. And I figured if it was this easy to make a model, one, we have a whole new way of monetizing data for the first time. you can just, like, take really valuable data and turn it into a service in $600. and that cost will probably go down over time.
Swyx [00:04:03]: When you - So sorry. when you say monetizing your data as, what eventually will become an MCP endpoint or as a training data for a model?
Anjney Midha [00:04:12]: Yeah, training data for a model.
Swyx [00:04:13]: Awesome.
Anjney Midha [00:04:13]: Like, an abstract way of saying like, “Hey, I have this data.”
Swyx [00:04:15]: Compress it into a model.
Anjney Midha [00:04:16]: Like, it makes sense for me in my product, but, like, I could repackage it in the form of a model and sell it. And so it’s just a whole new business model for the economy. It also, of course, provides, like, a way of following what Frontier Labs are doing, but in a way that, like, a single developer or a small team of developers can roll on their own. And so - Whenever you have an example of that, like a breakout app that’s doing really well, and then some framework for imitating it with - in your own flavor, you have an immediate ecosystem of, like an immediate ecosystem, like, should arise because there’s just a huge gap between the, like, decisions that the single company is making and all of the variations in those decisions that, like, a wider ecosystem can create themselves. And so then, you need a marketplace to, like, discover all of those, services and all of those products. There wasn’t any place on the internet that, like, was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why.
Swyx [00:05:29]: The closest would be Hugging Face.
Anjney Midha [00:05:30]: Hugging Face was the closest at the time, yeah.
Swyx [00:05:31]: They just started Hugging, like, a few years ago before that.
Anjney Midha [00:05:34]: Yeah, and Hugging Face also didn’t have the closed-source models.
Swyx [00:05:37]: Yeah.
Anjney Midha [00:05:38]: And they didn’- you couldn’t use the models at the time. and there wasn’t data about who was using them. There were, like, a bunch of differences between OpenRouter and Hugging Face, and those differences felt really critical to me, especially when I was just trying to learn about LLMs and, like, why people are choosing, like, Different little ones that are emerging over time.
Discord, Open Models, and the Origins of OpenRouter
Swyx [00:06:03]: Got it. And then, Ansh, no stranger to wanting more model diversity, at the time, you’re a couple of years into your Anthropic journey, which we covered in the previous podcast as well. What was your introduction to Alex?
Alex Atallah [00:06:16]: Well, the introduction was, I think, thirteen years before that.
Swyx [00:06:20]: Oh.
Alex Atallah [00:06:20]: But the OpenRouter handshake happened right over there, if you remember.
Anjney Midha [00:06:23]: Yeah.
Alex Atallah [00:06:24]: Which - So Alex and I, met, I believe as sophomores now, if I remember at the Stanford Review,
Anjney Midha [00:06:32]: That’s right
Alex Atallah [00:06:32]: Meeting for the first time.
Anjney Midha [00:06:33]: I think so, yeah.
Alex Atallah [00:06:35]: Yeah.
Anjney Midha [00:06:35]: Yeah.
Alex Atallah [00:06:35]: So Stanford Review was the libertarian newspaper on campus at Stanford that Peter Thiel started back in the day. And, whatever-- for whatever reason, I, Alex and I both showed up to one of the meetings, and I remember, the editor-chief was a mutual friend of ours. Lisa was really a really great editor-chief, where, part of an editor-chief’s job is to assign responsibilities to people and make sure the work gets done. and I, I may be misremembering the details, but I remember wanting to. It was surprising to me that at the time there was no dedicated technology section in the newspaper.
Alex Atallah [00:07:11]: You
Swyx [00:07:13]: Because it’s political, right?
Alex Atallah [00:07:14]: It is primarily
Swyx [00:07:14]: Like, it’s talking
Alex Atallah [00:07:15]: It originally started as like a
Anjney Midha [00:07:16]: Yes.
Swyx [00:07:17]: Yeah, states and all those things.
Alex Atallah [00:07:17]: Correct.
Swyx [00:07:18]: Yeah.
Alex Atallah [00:07:18]: But it, - To take us back in time, you may remember this, but, there was this technology, legislation that was being debated called, the Net Neutrality Act. And net neutrality is, like, inherently this political concept, right? It’s, it’s about the regulation of - internet broadband access. And so there was a community of us who were technologists, but also debating the politics of the technology. And I thought the Review would be a great place - to, like, write about that. And I was working on, I think, a net neutrality article, and I remember proposing, “Well, maybe we should start a technology section.” And Alex was one of the only people who said, “Yes, that would be cool.” And said. I forget whether we ended up writing stuff together, but - that’s when we first met,
Alex Atallah [00:08:03]: Was 2011 or twelve. I forget which year it was. It was one of those.
Anjney Midha [00:08:09]: Yeah.
Alex Atallah [00:08:09]: It was at Old Union, if I remember correctly.
Alex Atallah [00:08:11]: That’s where we used to meet. But, along the way, Alex and I have had a chance to, To hang out often. And probably the time when we had the most professional overlap was when I was running the platform at Discord, and it had become this explosive platform for crypto
Swyx [00:08:32]: Yeah
Alex Atallah [00:08:32]: And NFTs in the middle of the pandemic.
Swyx [00:08:35]: Which also, by the way, you were in charge of safety and security as well, right?
Alex Atallah [00:08:38]: I was the head of platform, which meant all of the crypto - the DAO and NFT launch security debugging fell on
Swyx [00:08:45]: And their phishing and.
Alex Atallah [00:08:47]: The phishing, the social engineering attacks, the katana DDoS that we were getting hit by. but it’s around the time I first started teaching security at scale at Stanford, CS 153. And Alex was on the, - at OpenSea at the time, and I was trying to figure out how we could defend against all these attacks that we were. Like, and at peak, I forget, if you remember how much NFT volume was running through
Swyx [00:09:10]: Discord
Alex Atallah [00:09:10]: Discord, but it was, like, a meaningful amount of, like, it was, like, several billion dollars in NFT volume of GMV, so to speak, were running through the platform, and it was all coming from OpenSea. It was these, like, buy, sell,
Swyx [00:09:20]: The
Alex Atallah [00:09:21]: Servers
Swyx [00:09:21]: The D in DAO is Discord.
Alex Atallah [00:09:25]: Yes. And so that’s when I think we had hung out professionally. But a year after that, OpenAI gave Discord early access to GPT. Sorry, three. No, it was five. Yeah, five, which is the RL version of three. And that’s around the time we made a Discord bot with, OpenAI for internal deployment, and that’s when I realized we would need. Like, since I was part of the deployment team.
Anjney Midha [00:09:50]: What was the use case?
Alex Atallah [00:09:51]: There were two that were. And there’s, there’s a post now called “Discord is Your Place for AI with Friends” that somebody sent me recently that I wrote, and published in twenty-three. But There were two use cases. One was Clyde, which was the - like, a party friend inside of Discord that could help you set up your Discord server and talk to you about onboarding and get your friends to hang out more. and then there was content moderation. And one of the realizations we had with content moderation was - it would refuse to moderate. Like, it would just refuse our prompts because the The training was. We were very early in the training era, and it would just. Our prompts would trigger it, its, like, guardrails. And we told OpenAI, “Hey, guys, we need access to the weights because if we’re gonna be doing content moderation at scale, we had 250 million monthly active users, we need more reliability that the model will do what we need it to.” And they said, “Well, sorry, guys, that’s not how this works. We’re a closed-source company.” And so that was my first realization that we needed open models, and the enterprises would need more control over capabilities, and then ultimately would need some control plane or management system to orchestrate these open models. But there weren’t no good - there were no good open alternatives until maybe
Alex Atallah [00:11:10]: Six months later when Llama came out. And six months after that, I led the series A into Mistral, which was started by Guillaume and the Llama team. And - That, - Around that time is when I remember hearing about Alex launching OpenRouter and going, “These worlds are gonna collide, and I don’t know when it’ll make sense to team up.” But Alex was so early and could see. I think he was totally right about this ecosystem starting with Llama that then needed, like, a, an easy layer to manage for, especially for. I was approaching it from the enterprise perspective because I had been that, like, the. As the VP of platform at Discord, it was my job to ensure that when we deployed models to, like, 250 million users, they did what we wanted them to. And that was very hard, because if you outsourced it to the labs and they controlled the guardrails and their guardrails are their safety policies. Forbid the model from responding to your prompts. That was quite catastrophic.
Swyx [00:12:05]: Yeah. But what, a moderation is the thing that they want to support. And obviously, beyond that, they would - OpenAI would work with you, presumably to give you a moderation endpoint, which they offer for free.
Alex Atallah [00:12:16]: It was an interesting use case, that - So they did give us a moderation endpoint. However, as you guys know, every Discord server is like a mini deployment of itself. And so the use case was instead of having human moderators that have to interpret the norms of the community, you just give the, - Often, like every, subreddit, Discord servers, public ones have their own rules that the user, the users create.
Swyx [00:12:41]: Oh, yeah. We run the LinkedIn Discord in. Yeah.
Alex Atallah [00:12:43]: And then humans used to read those norms and then enforce it every day manually, like observing each message in these communities. And these communities have like millions of users. So we had a 5,000+ person team globally in the, on the Discord content moderation team. These are outsourced contractors who had a really tough job. And so the idea was instead, if you could give the norms of that server To the LLM, then the LLM would do custom moderation for that server. It’s almost like a, like context moderation for that server. And many of those servers’ norms just violated OpenAI’s rules. And so - It was like we had our own custom eval. So each server had its own custom eval. But Discord-- at the time, OpenAI’s evals, we were all so
Alex Atallah [00:13:28]: Primitive in our thinking about how to deploy these LLMs that often the training prompts were super handed. It said, “Oh, anything about Harry Potter, anything that has trademarked content, don’- refuse.” And if it was a fan - Harry Potter fan community, this is a real use case, that had content moderation, the LLM would just refuse.
Swyx [00:13:48]: Yeah.
Alex Atallah [00:13:49]: And that was just not precise enough.
Anjney Midha [00:13:52]: Another one that we heard was like if someone was trying to write like a detective story, and there’s one chapter with a lot of violence, like maybe someone
Alex Atallah [00:14:01]: Right
Anjney Midha [00:14:01]: Like kills someone, the LLMs would just refuse to, like, help with that part of the story.
Alex Atallah [00:14:07]: Yeah.
Anjney Midha [00:14:07]: And then - like, we used to be like, okay, this is not like structurally inherent to LLMs. There must be, like, some choice out there so that I can, like, switch to another model, when I’m getting, like, a refusal or a bad result from the main one that I have. And that, like, tension also drove me for a marketplace.
Why “One Model Wins” Was the Wrong Bet
Swyx [00:14:28]: Yeah. I think that is well accepted now. What was it like back then when you were raising or, starting this? did people get it? what was the, some of the struggles? I like getting stories out of him about how other VCs don’t get it. So like anything you wanna, talk about, now - Let’s, let’s call it, that the early journey of OpenRouter is done, right? You can obviously talk about some of the early days stuff.
Anjney Midha [00:14:54]: Well, I was gonna say that, like, the biggest objection we got is big model win, which is - all of the
Swyx [00:15:03]: Scaling laws.
Anjney Midha [00:15:04]: Huh?
Swyx [00:15:04]: Scaling laws.
Anjney Midha [00:15:05]: Yeah, scaling laws, and natural network effects are just gonna accrue to one company, which will be - It’ll be a Google-style monopoly, just like how Google won the search market, by a large margin, and you’ll just be fighting for scraps at the end. That was probably the biggest objection we got. it is interesting that Google won the search engine race with such a huge margin. I think, like, had there been more interesting benchmarks or had, like, search engines been, - had people, like, seen them a little bit more like LLMs where they’re services that you can build companies on top of, that might not have been the case. but LLMs don’t merely have a user interface. They’re also, like, ways of building entirely new businesses. And, a Google-level monopoly would be like the Dutch East India Company times, quadrillion in magnitude because the whole economy ends up, like, depending on the one monopoly as well. So it didn’t seem like would be a really crazy outcome if that happened. And it’s also less likely because the economics of, like, creating good competitors are much, like, much more decentralizable.
Alex Atallah [00:16:25]: Everything Alex said is true, And I came at it from a completely different perspective, which
Swyx [00:16:31]: Yes, this is why we’re here.
Alex Atallah [00:16:32]: The scaling laws were never - In my mind, were always a feature, not a bug for why OpenRouter would be very valuable. Because, I was one of the first investors in Anthropic, and it was obvious to me that other researchers in our friends - I went to grad school for machine learning, and I just had a lot of friends in the ML community who it was very obvious to us that the bitter lesson holds. And so I was like, “Oh, fantastic. Now we have at least two proof points that compute scaling works.” It was OpenAI and Anthropic. and by the time I think we decided to team up on OpenRouter, I had already invested in Mistral and Black Forest Labs and Luma. So there was multiple model companies and teams that I was, working with.
Why Model Labs Struggle With Distribution
Swyx [00:17:14]: But you did other modalities, whereas this is literally
Alex Atallah [00:17:16]: Across different modalities, yes
Swyx [00:17:17]: Text.
Alex Atallah [00:17:18]: Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being created, and that this whole narrative of, like, Only one company will dominate like Google was, well, like maybe true, but one, I don’t believe that. But two, there was so much extraordinary innovation happening across several different research teams. But the shared problem I was noticing across all of them was often, the research teams were fantastic at figuring out how to reason about new capabilities. They think in terms of capabilities, but never - like, are not developer mindset-oriented. Like, what happens after the training is done and the checkpoint comes out? Like, you’d be shocked how, like, similar the early training teams at OpenAI, sorry, Anthropic, BFL, Mistral, were in their, like, default approach to. Taking their research out of the, lab and scaling their impact, which is often, oh, the checkpoint is done, put it out as an API, done, and then there’d be crickets. in the case of Claude, the first Claude checkpoint was done a year before they released it internally. And then ChatGPT came out, and we decided, okay, yes, it’s a good idea to release a Claude version externally.
Alex Atallah [00:18:34]: And they had no plan, like no plan for how to get developers to try it out. And so if you go to the Claude one blog post, you’ll notice there are, like, three developer examples for users of the API, and one is a Discord bot, and the second is Vivian, my wife’s startup called Juny Learning, ‘- And then there was, like, Notion, because these were all friends of, like, the Anthropic Because that’s how - like, last minute the planning was around, hey, once the model’s done training, how do you get it out to the world? There was no distribution platform that understood what developers needed, all the key management, provisioning, like, simple, like, endpoint management, versioning control. Like, all these things that the scientists and researchers go, “ that’s plumbing. I don’t really think about it.”
Swyx [00:19:15]: Implementation detail.
Alex Atallah [00:19:16]: Right. And instead, Alex came at it from that perspective. And so, it was so obvious to me that, like, every single lab I was funding would spend - like, literally sometimes billions of dollars into training, and then a checkpoint would be done, and there’d be crickets, like, during early access because they’re like, “Oh, that’s right.”
Alex Atallah [00:19:35]: It’s hard to use a checkpoint to make anything. You need a whole bunch of plumbing around it to make it usable by a developer. And so by the - I think - it was so obvious to me that a distribution platform like OpenRouter was critical to have in the ecosystem if we wanted there to be competition to Google. Like, unless-- ‘cause with Google, DeepMind is done training a new checkpoint, and then they push a button, and it gets blasted out across all their surfaces from Google Docs to,
Swyx [00:20:01]: Everywhere, even if I don’t want it.
Alex Atallah [00:20:02]: Everywhere. You wanna know about, like, on Android, like, overnight, they can deploy a new checkpoint to, like, a billion devices, right? And that invisible infra advantage, distribution advantage, most people don’t realize, but until OpenRouter showed up, - you had to think about all of that yourself as a model lab. And it was very daunting. at Anthropic, I think it took, well, more than twelve months to get to our first 10 million in revenue. And in contrast with Black Forest Labs, I remember the early days, you guys had a conversation with the BFL team, and, it was so simple for OpenRouter to say, “Oh, no problem. Like, the day you launch, we can send 1 million developers to you.” that was crazy. That was like a step function change in, like, an hour.
Swyx [00:20:46]: Is that a real number, a million?
Alex Atallah [00:20:47]: I,
Swyx [00:20:48]: Okay. All right.
Alex Atallah [00:20:48]: I think today it’s, like, 4 million. How many developers are on OpenRouter today?
Anjney Midha [00:20:52]: Over ten,
Alex Atallah [00:20:54]: Yeah.
Anjney Midha [00:20:54]: Over 10 million, but, like, it’s, it’s hard to, you
Alex Atallah [00:20:59]: I, yeah, I don’t know how to. Yeah.
Anjney Midha [00:21:00]: We do a lot of, like, account duping work, but, no
Alex Atallah [00:21:04]: If you could get 1,000 developers, just to put in context If you get 1,000 developers who try the model on day one after you release it and just, like, do inference and give you feedback, that’s a thousand
Anjney Midha [00:21:15]: That’s huge
Alex Atallah [00:21:16]: More developers than they knew how to get to on their own.
Swyx [00:21:19]: Well, BFL had a reputation, but yes.
Alex Atallah [00:21:21]: They had one in Stable Diffusion.
Swyx [00:21:22]: Yeah.
Alex Atallah [00:21:23]: And with Mistral, I don’t know if you guys remember, but the first checkpoint they released was, like, torrents. It was, like, torrent weights.
Swyx [00:21:31]: Yeah, they just put up a magnet link.
Alex Atallah [00:21:33]: Yeah, there was no API.
Anjney Midha [00:21:34]: Yeah.
Alex Atallah [00:21:34]: Because they didn’- they weren’t infra people.
Alex Atallah [00:21:37]: ? Like, it’s like, okay, download these weights, and you guys go figure out how to host it.
Swyx [00:21:39]: Well, he has a story on his side, yeah.
Anjney Midha [00:21:41]: Yeah, in addition to the, like, building a really good developer experience around it, the marketing that we do on, like, for different models is totally different and perceived totally differently
Alex Atallah [00:21:54]: Right
Anjney Midha [00:21:54]: From the marketing that a model lab does for itself.
Alex Atallah [00:21:56]: Yes, 1,000%.
Anjney Midha [00:21:57]: Right? We are like a, neutral layer looking at this market like it’s a big dark room with all the corners completely obscure to users, and users are walking into the room and, like, feeling around
Alex Atallah [00:22:09]: Yeah
Anjney Midha [00:22:09]: And trying to figure out what objects to grab off the tables and, like, build into, their companies. And it’s just an insane way of working. Like, models are not products where you can just enumerate all their features onto a web page. They’re all black boxes, including the open weight ones. So you need to, like, shine lights on all corners of this room, so that people can see what makes this model good, and you need the company shining that light to be a neutral third party, which is what we specialize in. So the, like. It’- In addition to developer experience, there’s also, like, a very important, like, marketing and product packaging component
Alex Atallah [00:22:50]: Yeah
Anjney Midha [00:22:50]: And a way of, like, routing and discovering models becomes, like, critical to your market as a provider or a model lab or a server tool and more in the future.
“Just a Wrapper”: Why VCs Misunderstood OpenRouter
Alex Atallah [00:23:03]: And this value, to your earlier point about how many VCs, like, just don’t. One of my biggest frustrations is that venture capitalists, many of them, like, just don’t have any operating experience in the field. so unlike a traditional investor who’s just maybe come up through the ranks as, like, a associate working on financial modeling or maybe hasn’t been a real operator in the field for, like, more than ten years, which is a big part of the industry now, I had just arrived at a16z, like, a year after running the platform. And so I knew what the challenges were of, like, building a real - great developer experience and like, being able to create a working piece of software with a model. And there were a few, I won’t name names, but there were investors who were looking at OpenRouter, and, felt at the time, like, when I would compare notes with people, that it was just, I quote unquote, “just a marketplace.”
Swyx [00:23:59]: Yeah, just a thin layer, just a
Alex Atallah [00:24:00]: Correct
Swyx [00:24:00]: Just
Alex Atallah [00:24:01]: A wrapper or whatever on other people’s APIs. And I was like, “You have no idea how strategic the value that OpenRouter has created by being able to orchestrate even three.” APIs in production. The amount of both engineering work and community design that goes into getting that live and running in production at the scale the OpenRouter team had started just doesn’t happen by default. And that was one of the things that stood out to me about Alex from the earliest days. Like, he just understood, like, - from a systems perspective, like, how do you get these flywheels going? Like, that stood out to me with OpenSea when we were working together on the NFT integration at Discord. Like, Alex had a level of community-- like, systems thinking on how you get these flywheels going that most scientists and machine learning people just don’t
Alex Atallah [00:24:48]: Think of. Like, we often think in terms of training.
Swyx [00:24:52]: It’s a linear stage.
Alex Atallah [00:24:53]: It’s this linear pipeline.
Swyx [00:24:53]: There’s no loop yet.
Alex Atallah [00:24:54]: Yeah. It wasn’t until much later that the modern context feedback loop cycle really got standardized in the industry. But at the time, if you remember, machine learning was like. Like, mostly we did a lot of ML, like, when I was in grad school on a laptop. So you just, like, download a dataset, ran some ablations, and you looked at the loss curves, and you’re like, “Great, I made AI.” And the idea that you have to, like, deploy those capabilities, collect feedback trajectories, then, like, put those into a continuous loop, like, came much later. And it was very counterintuitive to the - like, the traditional AI mindset. I do remember doing the investment phase for, OpenRouter, I just didn’t try and educate a bunch of other VCs on why it was not just a marketplace. I was like, “ what? I’m just gonna invest.”
Anjney Midha [00:25:41]: Yeah.
Alex Atallah [00:25:41]: And I’m going to, like, take the opportunity to partner with Alex, and if - no other VCs get it, that’s totally fine. ‘Cause at the time, - it was not obvious, I think, to several of the investors that, like, OpenRouter was not more than just a wrapper around APIs. And - that infuriated me. And I was like, “ what? I don’t have time to debate you. I’m - we’re gonna, we’re gonna invest.” And then I think, like, a month later, Matt Murphy marked it up by 10x. Like, - I think. I forget what the exact money was and so on, but, to his credit, Menlo Ventures realized, “Okay, there’s much more strategic value here as well.” Maybe you didn’t hear all these conversations behind the scenes But that frustrated me a lot. there’s a lot of this, like, opining about wrappers. and if you’re like, “Oh, an app is just a wrapper on a model,” then, like. And, OpenRouter is, like, this wrapper on top of other APIs, and this is the most stupid, reductive framework.
Alex Atallah [00:26:31]: And so it’s clearly somebody who has no experience deploying product at scale.
Swyx [00:26:34]: It’s the thing you dismiss other things with. Like, you’re a - everyone’s a wrapper on everything, right? Like, and there’s, there’s some Some wrappers have value.
Alex Atallah [00:26:40]: Investors are wrappers and LPs, right?
Alex Atallah [00:26:42]: Like venture capitalists. So, yeah, it’s all wrappers down, all down to bare metal, I guess, and like energy.
Swyx [00:26:46]: Yeah, there - When I started the whole AI engineer, I guess, the coining, in 2023, like, that was, like, the number one pushback is that this is no value. You should just train models.
Anjney Midha [00:26:56]: Right.
Swyx [00:26:57]: And, yeah, obviously this is, like. you guys are one of the testaments to the fact that you can build very valuable wrappers, but also very valuable model companies.
Alex Atallah [00:27:06]: It’s so, hard to be. Like, the day a model launches, the fact that you have an OpenRouter, endpoint for that model frequently at the top of Hacker News on day one, people don’t realize the amount of work that goes into accomplishing that. And OpenRouter used. Like, that would happen over and over again, and I remember going, “People have no idea how hard that is.”
Alex Atallah [00:27:30]: That’s not.
Swyx [00:27:31]: Yeah, we’ve covered some of the inference engineering that goes behind,
Alex Atallah [00:27:34]: Yes
Swyx [00:27:34]: Some of - with Base Ten and all those. Well, today you have, all those, like, cool code name things that people guess what Oxy Alpha is and all those things. But, like, I guess one of the things that you’re teasing is, how do you get that initial flywheel going, right? Because today you have your scale and your reputation, all these things, so obviously you - you’re driving immense distribution. But when you were early on, when it’s mostly
Bootstrapping OpenRouter Through Community
Alex Atallah [00:27:55]: The bootstrap, yeah.
Swyx [00:27:56]: Yeah.
Alex Atallah [00:27:56]: What was the bootstrap like?
Anjney Midha [00:27:58]: To bring it back to early Discord days, I think we, like, initially connected with. This is an OpenSea story, technically. But, and we initially connected when you were at Discord, and we talked about, like, - the Axie Infinity server.
Alex Atallah [00:28:13]: Oh, yes. Yes.
Anjney Midha [00:28:14]: This server was, like, the biggest server at the
Alex Atallah [00:28:17]: Yeah
Anjney Midha [00:28:17]: At Discord.
Alex Atallah [00:28:18]: That’s right.
Anjney Midha [00:28:19]: And you were like, constantly bumping up the
Alex Atallah [00:28:22]: The limits on the server. Oh, my God
Anjney Midha [00:28:24]: Of how many people could be in the server.
Swyx [00:28:24]: For those who don’t know, like, 10% of Philippines was Axie.
Alex Atallah [00:28:29]: Was on that server. That’s a big hit.
Swyx [00:28:31]: It was, like, a meaningful contributor to the GDP of the country.
Alex Atallah [00:28:33]: It was an NFT, like, crypto game, but it
Swyx [00:28:35]: It was like a Pokémon breeding thing.
Anjney Midha [00:28:36]: Yeah.
Alex Atallah [00:28:36]: Yeah. Similar. Yeah. There was battling, there was breeding, and then there was, like, a marketplace for trading.
Swyx [00:28:43]: Earn as well.
Alex Atallah [00:28:45]: Yeah, earn. And, like, the graphics were really cute and fun, and you like, you get emotional about your Axie that you make. So to, like, start a community like that, which we had to do many times at OpenSea with every early project, for us to create a marketplace for it, we need to make sure that the, like, the community wants it.
Anjney Midha [00:29:09]: Right.
Alex Atallah [00:29:09]: And it’s like building something that people want and going and telling them about it. Like, you can do that on a one basis, but there’s way higher leverage to do that in a community where everyone can talk to you at the same time. So we spent a lot of time, like, building things that the community really wanted. We did the same thing for OpenRouter. And, like, the Axie community was one of, like, a zillion communities we did that with. And Anj, like, saw us doing it and. ‘Cause you could just see people sharing OpenSea links constantly in that Discord. Like, users sharing links is a really clear indicator that, like, something important is going on. So we spent, a lot of time, like, first figuring out what the gap is in the technology that people care about. Like, what was the actual problem that needs to be solved? in early LLM days, it was, OpenAI refusing to finish the prompt or,
Anjney Midha [00:30:09]: Yeah
Alex Atallah [00:30:10]: To, like, complete the task. It was also.
Anjney Midha [00:30:13]: Inability to customize models. and so there are communities that, like are just completely blocked on that issue, and those are the communities that are most useful to learn about and dive into and explore.
Alex Atallah [00:30:28]: Something that really struck me at that time, - as I was just hearing your talk, I remember noting - you may not remember this, but we - we had these, like working, Zoom calls that we were doing a sprint around for, like this OpenSea integration with Discord. and, we’d, we’d - it was myself, my engineering team. I think you were there. And I remember, Alex, in the middle of one of those calls, just like there was like silence. we were all like, “Oh, yeah, this totally makes sense. Let’s do this.” And then there’s - every, like everybody aligned. And Alex was like, “No, this makes no sense to me.” And everyone’s - I remember going, “What? Like, it works. Like, you click on a link and this, then it bounces you out to, like, OpenSea.” And he was like, “It’s not a good user experience. Yeah, we should not do this.” And I remember going, he was the only one person out of all of us to raise his hand and go, yes, it made sense from a technical implementation perspective. Like, we were bouncing the user out into the, into OpenSea. And so it kinda checked the box of the product manager’s requirements on both sides. But Alex went one step further and was like, “ what would be better, guys? If we just embedded the experience right here inside of Discord so the link opened up as an embedded iframe, and you can just check out right there.”
Alex Atallah [00:31:47]: And not one person on the call, and there’s like seven of us who had met, like, week after week.
Swyx [00:31:52]: And it’s the guy who doesn’t work for Discord.
Alex Atallah [00:31:53]: And it’s the guy who doesn’t work for Discord.
Swyx [00:31:55]: Like, technically, you benefit if they bounce.
Alex Atallah [00:31:57]: Exactly. And that was, like, adversarial. To keep the user inside of Discord would be adversarial to OpenSea. And yet Alex put that user experience first. And I was like, “That’s special.”
Swyx [00:32:08]: Wow.
Alex Atallah [00:32:08]: Because it’s very hard to have somebody who’s technical like Alex and understands the developer flow, but also understands the best user experience and wants to prioritize that. And that’s two sides of the flywheel that if you can get spinning, like is often hard to stop. And you just reminded me, like that one was one of those moments where I go, I - I realized I gotta be better at user experience because I should have been the one who came up with that, and I didn’t. And I learned from you. And, I think that went into one of our case studies for the PM training program at Discord.
Swyx [00:32:34]: Whoa.
Alex Atallah [00:32:36]: I don’t know if it there is Because of
Swyx [00:32:38]: You need an Alex is the conclusion.
Alex Atallah [00:32:40]: Yeah. You need an Alex. And this is why I’m not, nobody should be surprised why Stripe decided like they had to buy OpenRouter because it’s a really rare combination of people who understand the machine learning community, the developer experience, and the user experience. And putting all that together has resulted in this extraordinary scale that very few other marketplaces have been able to achieve
Window AI, BYOM, and Finding the Right Form Factor
Swyx [00:33:02]: Yeah.
Alex Atallah [00:33:02]: Over the last, five years.
Swyx [00:33:04]: Yeah. Well, we should talk about the other reasons for acquisitions, which
Alex Atallah [00:33:07]: Yes, we should.
Swyx [00:33:07]: You’ve written about. I wanna proceed somewhat chronologically as well. So - there is a point that, one of the questions that, Dave from H of Zero sent in was, when did it - really started to work? And you brought up Mixtral. I don’t know if you wanna bring up that story.
Alex Atallah [00:33:22]: Oh, yeah.
Swyx [00:33:23]: Which obviously you overlap with, so.
Anjney Midha [00:33:26]: Yeah, the MoE was. I don’t know when. there’s no like one moment where I was like, “Oh, this is, officially starting to work.” It was
Swyx [00:33:36]: The moment where you had a Chrome extension, like, really super early on.
Anjney Midha [00:33:39]: Oh, yeah. But, well, - yeah. So before OpenRouter, I wanted to, like, explore a bring-your-own-model experiment. And,
Swyx [00:33:47]: Which anyone familiar with crypto is like, yeah, Phantom and all these things.
Anjney Midha [00:33:50]: Yeah. So it felt like doing a MetaMask analogy for AI would be a fun way of exploring that. And at the time, there were no AI apps. There were probably as many AI apps that were, like, hitting AI - like, hitting an LLM via an API call as there were, like, games just doing it in JavaScript. like there was a, there was a moment in time where it could have been the case that web apps call LLMs through the browser, like through some desktop
Alex Atallah [00:34:27]: Yes.
Anjney Midha [00:34:27]: Managed app that is controlled by the user. and of course, there are like, I think, many reasons that did not happen. But back when the days were that primordial, I built a Chrome extension called Window AI
Swyx [00:34:43]: With Plasmo.
Anjney Midha [00:34:44]: With Plasmo.
Swyx [00:34:45]: I had come across early on, and I was like, “Who’s gonna use this?” You did.
Anjney Midha [00:34:49]: Plasmo had a couple, like, I think Phantom was using it. there were some other, like real companies using it.
Alex Atallah [00:34:56]: It was like a shim.
Swyx [00:34:57]: React for Chrome extension. It compiles to all
Anjney Midha [00:35:00]: Yeah.
Alex Atallah [00:35:00]: I see.
Anjney Midha [00:35:00]: Like Next.js for Chrome extensions.
Swyx [00:35:01]: Next.js, Next.js.
Alex Atallah [00:35:02]: Okay.
Anjney Midha [00:35:03]: And yeah, built Window AI on top of it. The creator of Plasmo, like started contributing code to Window AI, in GitHub, and that turned out to be Louis Vicchi
Alex Atallah [00:35:15]: Oh, you’
Anjney Midha [00:35:15]: Who is the founder of OpenRouter.
Alex Atallah [00:35:17]: That’s right. You have told me this is how you met Louis. Yes.
Anjney Midha [00:35:19]: Yeah.
Alex Atallah [00:35:19]: Okay.
Anjney Midha [00:35:20]: So, that allowed users to like configure which model they wanted to use for a web page in their browser, and then, like the app would just call out to that model when it needed to do things. not the right form factor for LLMs, but, it’s like fun experiment. You learn a lot, and like I open sourced it. And the main learning is like, okay, this has to be an API, and it has to look a little bit - like, there has to be more of a developer experience here and more of a discovery experience as well. Like, I don’t know where to use these models, and a little Chrome extension is not gonna help me discover. It’s not enough real estate. I need more space. I need visuals. I need graphs. I need, examples. I need images. I need to, like, I need to be able to, like explore both as a human and as an agent.
Crypto, Midjourney, and the Early Generative AI Ecosystem
Alex Atallah [00:36:10]: Yeah.
Anjney Midha [00:36:10]: So that’s how OpenRouter came to be.
Alex Atallah [00:36:13]: A meta point that.
Alex Atallah [00:36:16]: I think is underappreciated, but Alex is reminding me, is that we were quite lucky that we were so. we were, like, adjacent to the crypto community in those days. Because in hindsight, crypto ended up being like a dress rehearsal for generative models, right? If you think about the Axie experience, Alex is totally right, there were not that many AI apps at the time. And while I was dealing-- my job was to be the head of platform at Discord, which meant to be a general purpose place for communities and friends to create-- for developers to create apps and bots and, other services that could be deployed across Discord. And while 80% of the attention at the time was being spent on crypto, because that’s where all the NFT volume was, there was, like, twenty percent of my time I was spending with a friend, who would get hotbot with me and ask me for. We would play Magic: The Gathering on weekends, and he was working on a little Discord bot that could take a text input and turn it into an image, and it was called Midjourney. You
Swyx [00:37:15]: Is that David?
Alex Atallah [00:37:15]: It was David Holz.
Alex Atallah [00:37:16]: He was a good friend. And David and I have both been failed ARVR founders, in the before that. And, I remember this. Midjourney was one of the fastest-growing communities we had after Axie Infinity started to peter off. And many of the, like, the abstractions and the infrastructure decisions we made to scale Axie happened just in time because they. Axie did this and then fell off a cliff. And then as Midjourney was taking off, we, like, explicitly decided to help David make the server, the Midjourney server, as the primary place for interaction with the model, because it was very hard for people to understand how to use the model if they couldn’t see other people using it and copy them. And so the single-player Midjourney web app on its own, like midjourney.com, had, like, terrible retention because people would show up, they’d see this empty field. It’s like E 2, and they would type in, like, cat or dog. And it was, like, paralyzing for them to have this blank canvas that they had to fill because they’d never used an AI model before. But instead, in a Discord server, you could see other people using it and riff off of their prompt, and the engagement was off the charts. And so scaling, Midjourney from zero to, like, 10 million monthly actives was a much smoother approach Axie Infinity. And so,
Swyx [00:38:29]: Don’t forget the best of four pictures, and you choose one.
Alex Atallah [00:38:31]: The best, yeah, and then the other, we
Swyx [00:38:32]: Which is the feedback loop.
Alex Atallah [00:38:33]: The RLHF feedback loop, which, by the way, separately, like, Tom Brown, David and I used to play Magic: The Gathering on weekends. And so, like, it was one group of friends would hang out, and we’d. Like, these concepts were all being discussed all the time. But, there was.
Alex Atallah [00:38:47]: I think there were few of us who bridged both the crypto worlds and the AI worlds. And compared to crypto, where it was - the question was always, what’s the use case, for this technology? There was never any need to ask that for AI because it’s, like, the use case was so visceral. It was like, I can create now anything at - I can imagine. I can write novels, I can code. And the infrastructure that those of us who believed in the distributed systems, like, value of crypto, like the censorship resistance part, found this use case that was explosive. And I think between Midjourney, the, Claude was a Discord bot launch, that we were using internally as an LLM. ElevenLabs had a TTS model that we had on Discord as well. Like, Discord became this petri dish for, like, early apps to innovate. And I don’t think it’s a coincidence that they found a home there before OpenRouter gave the world, like, a public home store or, like, a, storefront. Discord was this, like, almost petri dish storefront that - had, like, piggybacked on the infra we’d built for crypto communities. And then I think Alex was one of the first people to realize, wait a minute, like, these apps need their own home, on the internet. And then OpenRouter, to me, was a continuation of that community’s needs. And of course, there was the crazy distribution that you enabled for a lot of these developers.
Why OpenRouter Couldn’t Just Live Inside Discord
Swyx [00:40:07]: So then my question is, how come you were. My perception is OpenRouter is not that Discord-centric, right? You have a Discord.
Anjney Midha [00:40:14]: Yeah.
Swyx [00:40:14]: And you use it to engage your community, but it’s not like Midjourney where, like, no, that is like the primary way people experience OpenRouter.
Anjney Midha [00:40:21]: Yeah, Midjourney, like, it really helps to see visually really quickly how people are using the model and how to prompt it.
Swyx [00:40:29]: Yeah.
Anjney Midha [00:40:29]: And I think that is partly why the server was so critical. It’s like it is the user experience. It adds a ton.
Swyx [00:40:36]: Yes.
Anjney Midha [00:40:37]: And you can go the whole mile with just, like, prompting via Midjourney, like, the, via the Midjourney Discord server, getting your images and then sharing them and having fun. For OpenRouter, for LLMs, like, you need a lot of user experience around LLMs to make them, like, really usable.
Swyx [00:40:54]: Charge point.
Anjney Midha [00:40:55]: And yeah.
Anjney Midha [00:40:57]: The, like, seeing the examples of other people is also not as useful because it’s a lot of stuff to read. It takes a long time.
Swyx [00:41:03]: Yeah.
Anjney Midha [00:41:04]: You need, like, based integration. Not possible to do in a Discord server. You need, Or technic- it’s possible. I shouldn’t say that. It’s just not a great developer experience. you need, like, - you need governance for. At the point where you got based integration, now you need governance for managing the LLMs that have access to it, the data policies, which teams. All that stuff needs a lot more than a Discord server can provide. So it’s just
Swyx [00:41:30]: Yeah
Anjney Midha [00:41:30]: It’s not the right.
Alex Atallah [00:41:32]: Well, in addition, you’re not wrong, but also there’s the very important distinction that, Midjourney was an end user application.
Swyx [00:41:40]: Right.
Alex Atallah [00:41:40]: And, that’s why Discord, which has 250 million monthly end consumers, made, it made sense for Discord to be a host for that application experience. What I knew was gonna happen soon after Midjourney found explosive product-market fit, because we. I think when Midjourney launched, from launch to $100 million revenue run rate, it was less than eight months. And shortly thereafter, Stable Diffusion launched. And, all of us used to hang out in the Discord server. There, I think it was the,
Swyx [00:42:13]: The Stability Discord?
Alex Atallah [00:42:14]: It was the
Swyx [00:42:16]: Yeah, LAION.
Alex Atallah [00:42:16]: Yeah, the LAION Discord server.
Swyx [00:42:17]: The image community that spawned Stable Diffusion.
Alex Atallah [00:42:19]: The image community. Yeah. And so when Stable Diffusion came out, I realized- Oh, now other people can build their own Midjourney.
Alex Atallah [00:42:27]: Because until then, Midjourney did not have an API, so they were a stack company, right? They were training their own models, and they were deploying them as an application. But if you wanted to build your own Midjourney, there was no API of that quality. and I think E two was still quite primitive. Like, Midjourney had great quality. And then when Stable Diffusion came out, suddenly there was this new person who - there was - this new capability in the world, which is a developer could create their own Midjourney. And that, I think, created the need for something like OpenRouter, because then you need an API to. If you - if you had the creativity of David Holz and you had Stable Diffusion as the model and you wanted to put these things together, how could you do that without having to figure out how to host the weights? And what OpenRouter, - the shape of OpenRouter enabled is that. Right? When you have open model alternatives to closed applications, OpenRouter’s value in the world becomes extraordinary because now any developer can just show up and use the
Stable Diffusion and the Need for a Model API Layer
Swyx [00:43:20]: You just love model diversity.
Anjney Midha [00:43:21]: Did you just say the shape of OpenRouter?
Alex Atallah [00:43:23]: Oh, no.
Anjney Midha [00:43:25]: Were you in cloud? What is this the real Han?
Alex Atallah [00:43:26]: I’ve been, I’ve been - I’m, I’m misaligned now. I’ve been overtrained. I’ve been using Cloud way too much, haven’t I?
Swyx [00:43:34]: Claude-ish is what people would say.
Alex Atallah [00:43:35]: Claude-ish. Oh, God, I gotta untrain myself.
Swyx [00:43:38]: Okay. - And I just wanna cap off the Mistral side. my TLDR is there was a Mistral price war, is what they called it, right? Like, round about NeurIPS is twenty-three or twenty-four.
Mistral and the Birth of the Inference Marketplace
Anjney Midha [00:43:47]: Yes. December
Swyx [00:43:48]: They launched, the Mistral 8x7B, and like the price went down like 80%.
Anjney Midha [00:43:54]: Yeah.
Swyx [00:43:54]: To me, that’s very positive because it’s like the first, like, real competition to host Mistral. Is there more?
Anjney Midha [00:44:01]: Yeah, that was. I’m, like, trying to remember it, all the things that happened. It. Like, we saw that model come out and immediately saw people say that it was the best model in the world.
Alex Atallah [00:44:15]: Yes.
Anjney Midha [00:44:15]: Like, this was, to my knowledge, the first time an open weights model was called that in real seriousness.
Swyx [00:44:22]: It’s hype, right? Is it?
Anjney Midha [00:44:25]: It was hype. It was hype. It was also, like, hype from AI influencers at the time. And there were many examples where it was, like, outperforming four. So people really wanted to try it out and see, is this gonna be true for me too? And if so, at what price? And, the, like, inference landscape was really messy.
Alex Atallah [00:44:49]: Yes.
Anjney Midha [00:44:50]: We cleaned it up. - it allowed, like, providers to compete on price, so we could give you just the best price in one spot. And so it was, I think, the first clear example of, like, a provider marketplace working in a way that adds value to end developers.
Alex Atallah [00:45:08]: Sean, you may not remember this, but I think we met for the first time a few days after Mistral came out at NeurIPS
Anjney Midha [00:45:15]: Yeah.
Alex Atallah [00:45:15]: At a luncheon.
Swyx [00:45:16]: Yeah. That’s where I also met BFL as well. Yeah.
Alex Atallah [00:45:18]: And Guillaume was there.
Swyx [00:45:19]: Yeah.
Anjney Midha [00:45:19]: I was at NeurIPS at that time.
Alex Atallah [00:45:20]: You were there too. And, we had just announced the Mistral investment, and I remember Guillaume was over there, and I remember turning to Guillaume and asking him, Like, “Is it is all the. Like, how are you feeling after the launch of Mistral and seven B?” And, him in his typical French fashion was like, “ it’s a, it’s an okay model. It’s not that good.” And I was like. It was so, in contrast. But I remember him also saying that part of the reason he felt a lot of people Thought that it was better than four was because of the speed. - it was an MoE model that they had, like, absolutely figured out how to make super efficient. It was on the Pareto frontier. And this is an important thing about LLMs, right? Sometimes when they’re faster, you think they’re smarter, even though, like, if you did, N of, these common, like, evals that are - you do seven tries, and I don’t remember. I think we should go back and figure out what the data says, but I wouldn’t be surprised if it turns out, oh, on an N of seven attempts, four was smarter on evals, but the perception of on, like, or correctness would be smarter or more accurate. But, people, like, from a human preference perspective felt that it was faster because it - or smarter because it’s so fast.
Swyx [00:46:36]: Yeah. And most queries do not take that level
Alex Atallah [00:46:39]: Don’t take that. That’s true.
Swyx [00:46:40]: Right? So this is the start of humans as router
Alex Atallah [00:46:42]: Yes.
Swyx [00:46:42]: Which then eventually becomes OpenRouter as router of like the
Alex Atallah [00:46:45]: Oh, that’s interesting way to think about it. Yeah.
Swyx [00:46:47]: Like, because humans are the routing mechanism. Like, I will ask the fast model first, and then if, like, oh, not good enough, I’m gonna upgrade manually.
Alex Atallah [00:46:52]: Yes.
Swyx [00:46:53]: But then he’s gonna auto it.
Alex Atallah [00:46:54]: I didn’t, I hadn’t thought of it that way, but that makes sense.
Swyx [00:46:57]: Which then there’s, there’s a lot more techniques, like fusion. Fusion is the thing that we should talk about. Before I move on to those things, I just want to close off the early years. one thing that I observe, which you are also an investor in Arena.
OpenRouter vs. LM Arena
Alex Atallah [00:47:10]: Right.
Swyx [00:47:10]: And we talked about Midjourney having that feedback loop of, A, B, C, D, and choosing that very. being very important. And you understand the flywheel. So how come you didn’t build Arena, and how come Arena didn’t build OpenRouter?
Anjney Midha [00:47:23]: Well, Arena started before OpenRouter, right?
Swyx [00:47:27]: They had the school project
Anjney Midha [00:47:29]: Yeah, LM
Swyx [00:47:29]: And then it became a company.
Anjney Midha [00:47:31]: LM Arena, yeah.
Swyx [00:47:32]: So, but, and I know you had some Arena experiences, like the up comparison type things.
Anjney Midha [00:47:37]: Yeah.
Swyx [00:47:37]: But you never really went as hard as Arena did.
Swyx [00:47:40]: And,
Anjney Midha [00:47:40]: In doing up experiences?
Swyx [00:47:42]: Yes. And LM Arena did have a router project based on LM Arena ELOs, which they never commercialized.
Anjney Midha [00:47:48]: It’s hard to do a company that does both because one company is taking data and selling it, and the other company really can’t by default. So, I think there is, like, a branding reason that there are two companies here. like, when you set up OpenRouter, there’s no training, there are no prompts, right, aside from what your provider policy set. Like, OpenRou- like, OpenRouter can’t see your prompts or completions. If you want to see that as an org, you have to opt into it and enable it. And so we’re, like, pretty conservative and careful about data policy and security. And privacy. And LM Arena is like, their business model is like oriented around the labs and,
Swyx [00:48:34]: Because they give it for free, right? You don’t give it for free to give it for free.
Anjney Midha [00:48:37]: Yeah.
Anjney Midha [00:48:38]: But we do give some. We like have free endpoints too, but like those free endpoints, we, I think we’re not collecting any prompts. We’re not like monetizing the data unless you, opt into it for some reason.
Alex Atallah [00:48:48]: This comparison. you’re not the first person to ask me this, and Alex knows this, but I was the interim, like the founder, like first CEO of Arena for the first five months when, and we were helping Anastasios and Waylin spin out of Berkeley. And, I did invest in that before, OpenRouter, but it was very strange to me the comparisons that outside, folks would make between the two projects because the missions were completely different. The founding entity for Arena, we called it the AI Reliability Institute because it was there as an eval service. Like the data, so to speak, that they were originally, offering the labs was how do you make the evaluation of models more reliable than like the state of the art at the time, which was like really just finger in the wind.
Alex Atallah [00:49:38]: That’s what Anastasios and Waylin’s PhD work was as scientists at Berkeley, was on statistical methodologies for correcting, eval estimates, based on like intrinsic biases and how you collected the data.
Swyx [00:49:54]: Yes.
Alex Atallah [00:49:54]: And
Swyx [00:49:54]: Style control.
Alex Atallah [00:49:55]: Style control and stuff like that. And which is very much like a, hey, how. If you’re a scientist and you’re trying to. the highest expectation customer for Arena was always like a training and, like a researcher at a lab. Whereas the highest expectation customer from my perspective that Alex like really understood and was the mission was to serve was like a developer, right? Who then takes the result of the research and then produces an application that’s deployed to the world. It was a completely different problem and person that these two teams were focused on. And so from the outside in. I don’t know if you remember this, but I have a distinct memory of a few weeks before we did the term sheet, together for OpenRouter, I’d given you a call because we were trying to get a pooled data set together from OpenRouter and from Arena to, create like an open source repository of prompts. these projects were so different in their goals that it was totally normal to me to be like, “Oh, yeah, let’s call Alex and see if he’d want to team up on pooling data,” because they’re so different. We need. We don’t have that data at all. We. Like, we didn’t have API prompts. We didn’t, we didn’t have like what developers want to do with the models, which is very different from what researchers inside a model lab want to do before releasing the model.
Swyx [00:51:15]: Yeah.
Alex Atallah [00:51:15]: Does that make sense? And so to this day, I think you see that this difference, even though at a 30,000-foot level you could. I guess you could conclude that Arena and OpenRouter are adjacent, but, the roadmaps, the missions and so on at the time at least were like in very different directions.
Swyx [00:51:36]: That ideal customer, I get. I totally get that.
Alex Atallah [00:51:39]: Yes.
Swyx [00:51:39]: As a founder, I want to own everything, right?
Alex Atallah [00:51:41]: That’s possible.
Swyx [00:51:42]: Like this is clearly an adjacency that I’m like gonna explore that.
Anjney Midha [00:51:45]: Own everything meaning like you don’t know what to do yet, so you wanna like make sure you catch PM
Focus, Anthropic, and Roads Not Taken
Alex Atallah [00:51:51]: No, I think what he
Anjney Midha [00:51:52]: As quickly as possible.
Alex Atallah [00:51:53]: You want to own the entire infrastructure space, and so you expand to whatever demand you can capture.
Swyx [00:51:58]: You want to have a play in each end.
Alex Atallah [00:51:59]: Yeah, I think that’s, that’s hard, in reality, because serving multiple customers is difficult.
Swyx [00:52:05]: Clearly, this is the one focus, right?
Alex Atallah [00:52:08]: Yeah.
Anjney Midha [00:52:08]: Yeah. I still think even in the age of AI, like focus is,
Alex Atallah [00:52:12]: Is critical
Anjney Midha [00:52:13]: Underrated and critical, not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is.
Alex Atallah [00:52:22]: One thousand percent.
Anjney Midha [00:52:23]: The world can map like, “Oh, I have this issue. Which brand out there is going to help me with that issue? This is the brand that’s known for that focus.”
Alex Atallah [00:52:31]: Yes.
Anjney Midha [00:52:32]: So like if I want real attention on this issue, like this really matters to me, I should go with the brand that cares the most about it.
Alex Atallah [00:52:39]: To underscore Alex’s point about how important focus is, in the early days of Anthropic, it was not easy to. Like people think that the early days of Anthropic were like super easy because they were on their 3 guys who left, but it was very competitive. The company was starting 10 billion dollars behind OpenAI, right? And so to get to the frontier, like the big question was, what do we want to be known for? What’s the mission? And the mission was AGI pair programming. And so to the, exclusion of all kinds of other things that were really shiny at the time, like image models and video models that were getting lots of, momentum, the Anthropic team was like, “We just got to focus on coding.” Like that is the core capability that we’re focused. And today you can see the results, right? It’s a trillion-dollar company within five years. And that focus, I think, like the high. The focus on who your highest expectation customer is and how you exceed their expectations, because exceeding anyone’s expectations is hard, and doing it for multiple like customers is so even more difficult, is part of the reason why OpenRouter succeeded and Anthropic as well.
Anjney Midha [00:53:39]: Was the focus on coding that early, though, or did it come later?
Alex Atallah [00:53:42]: Literally from day one it was AI pair programming is. Responsibly commercialize an AI pair programmer was the seed memo. That was when I invested, right? We like refined that memo a lot. Well, you got to ask Dario and Tom for permission on that.
Alex Atallah [00:53:57]: But it’s an extraordinary piece of writing that they had put together. And AI, commercializing it. Responsibly commercializing an AI pair program was the mission, from day one. And I would say there were maybe like a couple moments in the company’s history where like they did experiments to see if like little detours made sense, like a general chatbot, like Claude.ai when ChatGPT was really taking off. But, at the end of the day, but especially once, they got their like significant training compute online, I think like the. All the main evals at the company, for example, have always Coding evals, long horizon agentic programming. from day one, that was always the plan.
Anjney Midha [00:54:34]: Because when, like, Claude Instant came out and Claude 2 came
Alex Atallah [00:54:38]: Yes
Anjney Midha [00:54:39]: I remember the marketing mostly being focused on pros. Like, this
Alex Atallah [00:54:43]: Yeah
Anjney Midha [00:54:43]: Could write better
Swyx [00:54:44]: Yeah Long context. It was the first of its kind.
Anjney Midha [00:54:47]: Long context,
Swyx [00:54:49]: This directly affected me ‘cause I built something on that. Yeah.
Alex Atallah [00:54:51]: What did you make?
Swyx [00:54:52]: A small developer, which was my Devin before Devin.
Alex Atallah [00:54:54]: Oh, yeah. Yes.
Anjney Midha [00:54:55]: Yes.
Alex Atallah [00:54:55]: Small.
Swyx [00:54:56]: Yes. and, so I think, like, there’s, there’s all that really, like, good, like, focus is another thing - That is a question that people do wanna ask. you could have built any other things. Like, and obviously OpenRouter was working. were there other ideas that you wanted to pursue that you turned down? just the paths, roads not taken.
Anjney Midha [00:55:16]: We made a couple prototypes for things that we didn’t launch. One was a tuning model as a service.
Swyx [00:55:23]: Yeah. Lots of that with OpenPipe and, all those things.
Anjney Midha [00:55:25]: But it - It was in a very consumery form factor, where you would give us a YouTube video or two or three. We would then extract all the transcripts from it and try to tune a model to talk like the person in the YouTube
Alex Atallah [00:55:40]: Yeah
Anjney Midha [00:55:40]: Or the people in the videos that you sent. So, like, a really easy way of creating a tuned model based on, like, some videos that you like.
Alex Atallah [00:55:48]: That would be so useful.
Anjney Midha [00:55:50]: We,
Alex Atallah [00:55:51]: No
Anjney Midha [00:55:51]: We made it too. It was
Alex Atallah [00:55:53]: You don’t think so?
Anjney Midha [00:55:54]: It was, it
Alex Atallah [00:55:55]: And nobody used it?
Anjney Midha [00:55:55]: It - We didn’t like, test it with that many people because the model marketplace was our main focus, and it was, like, growing, and we were building more conviction in it over time.
Swyx [00:56:09]: Just, you
Alex Atallah [00:56:10]: Yeah. Why,
Swyx [00:56:10]: As a creator
Alex Atallah [00:56:11]: Yes. I’m a creator.
Swyx [00:56:11]: Have you been pitched many, like, - I have five hundred hours of recorded voice of myself.
Alex Atallah [00:56:17]: Right.
Swyx [00:56:17]: Make a thing of you, charge access to it. it works for OnlyFans, doesn’t work for
Alex Atallah [00:56:23]: I see
Swyx [00:56:23]: As regular people. I think - this is mostly, - It’s just a glorified RAG bot.
Alex Atallah [00:56:28]: Right.
Swyx [00:56:29]: Whether it’s in the weights or it’s outside the weights, doesn’t really matter. You’re just doing RAG on the videos, and people ultimately always just wanna find the source video, that directly answers it.
Alex Atallah [00:56:36]: Oh. my use case was mostly to practice - - with myself ‘cause I often like to see what. Like, the way I practice for a job interview or if I’m hiring a candidate or public speaking or whatever is I wish there was, like, a good
Swyx [00:56:48]: Yeah
Alex Atallah [00:56:48]: That I could, like, critique ‘cause it’s kinda hard to pull yourself out. I would never get. I would never offer it to other people as a service.
Swyx [00:56:54]: I wish there were, like, pick your top five mentors that, then talk to them instead of talking to yourself.
Alex Atallah [00:56:57]: That’d be cool too, yeah.
Anjney Midha [00:56:58]: That was, that’
Swyx [00:56:59]: That’s the creator AI. That’s a replica.
Anjney Midha [00:57:01]: And that was the use case we were aiming at.
Alex Atallah [00:57:02]: I see.
Anjney Midha [00:57:03]: Is like, you wanna create an experience
Swyx [00:57:06]: Like AI Steve Jobs and.
Anjney Midha [00:57:07]: And AI Steve Jobs was the initial use case.
Alex Atallah [00:57:11]: That’s a,
Anjney Midha [00:57:12]: Even though it’s not allowed.
Alex Atallah [00:57:14]: That’s a, that’s a common prototype, yeah.
Swyx [00:57:15]: Talking about adjacencies, tuning as a service, as part of the router service is something that I would typically think about as well, right? Like, why don’t you do that? ‘Cause if people are running already their inference through you, store everything, log everything, tune to a smaller model that is cheaper, faster, all these things that’s within your control, right? you didn’t do that, but, like, other people would have pitched that in the general state of a infra startup.
Anjney Midha [00:57:37]: Yeah. Yeah.
Alex Atallah [00:57:37]: I think you were just maybe a little bit early ‘cause today that’s an extraordinarily growing segment. Like, from Mistral, where they do a lot of enterprise deployments
Fine-Tuning as a Service and Infrastructure Adjacencies
Anjney Midha [00:57:44]: Right
Alex Atallah [00:57:44]: And stuff and tuning as, custom models for ASML or whatever. And often
Swyx [00:57:48]: But not as a router. They’re, they’re just like, “I come to you because I like your Mistral models. I want custom Mistral model,” right? It is not, “I want, to run all my OpenAI prompts, - store all my results, and then just move off of OpenAI.” Right? They’re not doing that.
Alex Atallah [00:58:01]: As a, as like a way to export off of dependency on a Frontier lab, I have not seen that yet. Yeah.
Swyx [00:58:08]: Right.
Alex Atallah [00:58:08]: Which was your vision.
Swyx [00:58:09]: Is efficient to do.
Anjney Midha [00:58:10]: We decided. Really, we, like, leaned into our focus and figured that, like, there aren’t. Like, we just saw the ecosystem develop over time. All these inference providers that do wanna help companies do that, - Like, it makes sense for us to partner with them and to, like, give users lots of choice and to, like, figure out what makes them, what gives them competitive advantages. It’s, it’s a whole new business and there’s, there’s value in being a neutral marketplace that just like, works with those companies.
Alex Atallah [00:58:45]: Could you share a little bit, to Sean’s point, like, how you prioritized. What are some ways you prioritize features? ‘Cause you’ve always done it so elegantly that I never. it just happens, and you make all the right decisions that always have product-market fit from the outside looking in. But consistently, you seem to have prioritized, a lot of hit features that worked. And maybe I have a sample set bias or whatever, but Sean’s question
Swyx [00:59:06]: Can you list what you think hit features worked?
Alex Atallah [00:59:09]: Oh, the leaderboards.
Swyx [00:59:10]: Leaderboard, okay.
Alex Atallah [00:59:10]: Yeah. like, from day
Swyx [00:59:13]: That’s charting, right? That’s the feedback loop.
Alex Atallah [00:59:14]: Charting, BYOK.
Swyx [00:59:15]: But, like, he had, like, ins. he had, like, And I think there was a whole thing I wanna get into about, like, completions versus
How OpenRouter Prioritizes Product
Alex Atallah [00:59:22]: Yes.
Swyx [00:59:23]: Check completions versus completions. And then also, let’s call it, like, the rise of the reasoning models and how you deal with those, multimodality, all those things, right?
Alex Atallah [00:59:31]: Yeah. BYOK.
Swyx [00:59:32]: BYOK, yeah.
Alex Atallah [00:59:32]: That was a huge one.
Anjney Midha [00:59:34]: There’s one I. Like, I think it was in early 2024, very early 2024, we thought it might be interesting to fuse the results of multiple models together, and we launched a prototype called MOM, Mixture of Models, that let you, like, pick a couple models, or we’d pick them for you, and then it would fuse the results together at the end, and it would show you all the intermediate results in this, like, big Kanban looking product.
Mixture of Models and Model Fusion
Swyx [01:00:05]: What does the fusion at the end, another model?
Anjney Midha [01:00:07]: Another model. The,
Swyx [01:00:08]: The smartest of
Anjney Midha [01:00:09]: The smartest
Swyx [01:00:10]: Of the set
Anjney Midha [01:00:10]: Of the three, of the set.
Swyx [01:00:12]: Okay. So this is like a council idea?
Anjney Midha [01:00:13]: Yeah. It was a model. It was like a very early LLM council.
Alex Atallah [01:00:16]: This is a agent swarm as, like, they would call it at one of the Frontier Labs, in the early days?
Anjney Midha [01:00:23]: Yeah, like some of those ideas are, like, going the right direction, but the devil’s in the details.
Swyx [01:00:27]: Yeah.
Anjney Midha [01:00:27]: There’s a lot of, like, product refinement needed to make them really work. they take your focus away
Swyx [01:00:34]: Right
Anjney Midha [01:00:34]: Whatever else you have going on. And there’s a lot of, like, community building and learning that you need to do. And the technology might be too early. So there are - like, all kinds of reasons they might go wrong. And in our case, the technology was a little too early. In other words, the fused result was a little bit
Swyx [01:00:53]: Right. Like a Frankenstein
Anjney Midha [01:00:54]: Sometimes the same as the best model that was being used to fuse because the best model was so far ahead of options two and three at the time. over time, the top three or four LLMs have gotten closer together, still neurodivergent, but, like, all capable of inserting, like, pretty interesting ideas. Like, RL has like, expanded the surface area of creativity for machine learning researchers within each lab, and so they can, diversify the reasoning power of different models more effectively. At least that’s my theory for
Swyx [01:01:29]: Yeah
Anjney Midha [01:01:30]: Fusion - it, like, works better than it used to, but early twenty-twenty-four. And, so the technology was a little bit too primitive. The form factor was not right, and so we would have had to go through a couple more iterations. And so we decided to just delete all the code. And, then years later, middle of twenty-twenty-six, or early twenty-twenty-six, we’re like, “Let’s bring it back.” Like, the research is looking kinda promising for fusion. The models now have, like, two, three, four top frontier models that are all really good and, like, I’m, I’m frequently trying to, like, consult multiple models to get the best results. Like, and then I ran a little personal experiment where I was like, “I’m gonna, like, do a, an architecture plan for a code change. I’m gonna give it to all the models. I’m gonna fuse the result, and then I’m gonna ask all the models if the fused result is better than the individual result each model came up with.” And they all said yes, that the fused result was better. And this happened a couple times, and I was like, “Okay, spot check, pretty good. We should, like, benchmark this.” And that’s how we built fusion.
Revisiting Fusion as Frontier Models Converge
Swyx [01:02:40]: Yeah. And it came on your Fable, so you were like, “This is Fable level.”
Anjney Midha [01:02:43]: Yeah.
Swyx [01:02:44]: Let’s start leading up to this year, which we haven’t gone to this year. can you mark out the main milestones in the journey? I think, it seems like your promise, was, routing. You decided the business model very early.
Swyx [01:02:59]: You take a cut. And, like, what are the major milestones that, inflect the growth, right? Like, you’re, you’re growing, like, 9% week on week now? Is this the official number?
Anjney Midha [01:03:10]: In terms of token volume, I think that sounds about right, yeah.
Swyx [01:03:13]: Yeah. So just, like, can you mark out, like, the brief history of OpenRouter up to, the acquisition? Let’s, let’s call we’re, we’re just, we’re just, talking about, people are, - you have a your birth moment with, the Mistral stuff where people are really competing. You have your state of AI thing where,
Anjney Midha [01:03:32]: Yeah.
Swyx [01:03:32]: It’s very cute. You have a hundred trillion tokens, ha, ‘cause now you’re doing ten a week, .
Anjney Midha [01:03:39]: Yeah. We’re doing ten a day.
OpenRouter’s Growth Inflections
Swyx [01:03:41]: Ten a day now?
Anjney Midha [01:03:42]: Yeah. More.
Swyx [01:03:43]: So yeah, you do this in ten days.
Swyx [01:03:45]: Like, what are the major end points there? I just wanna. Like, there’s a smooth curve, but, like, you feel the inflections.
Anjney Midha [01:03:51]: A lot of this is oriented around model launches. we had, a huge focus on pros all the way up through May of twenty-twenty-four, because coding was just not there, and no apps were able to build much on top of it. So, a diversity in models, but not a wide diversity and not a wide diversity in use cases. Dream Tavern was one of our top apps at the time. The creator of Dream Tavern now runs product at Cognition, Devon. - Then - In the middle of twenty-twenty-four, we saw Claude 3.5 Sonnet. That came out, incredible leap forward in coding, and we saw the dynamics of, like, apps building on top of us change. we saw a huge surge in volume in, like, users, using OpenRouter. And this is when I think people started to look at the, like, money that they were spending and get a little bit like, “Whoa, what’s going on? I might need to, like, think about, like, more efficient but equivalent models.” And shortly after that, I think it was after Sonnet three five, Mixtral 8x7B came out, and everyone was like, “What? This is the model.” Like, the OpenWeights community delivered. And so it was really good timing from Mistral.
Swyx [01:05:17]: All of Anja’s portcos are just helping you out.
Alex Atallah [01:05:21]: It takes an ecosystem to grow an OpenRouter?
Anjney Midha [01:05:24]: Yeah, that was the. Yeah, it was. It like, it was the, like, this early ecosystem, it was like a swing action where, like, model labs would come up with some frontier innovation. Like, usage would surge. Then users, look at their invoices 30 days later and like, “Whoa, what’s going on here?” And then OpenWeight models would deliver, like, a, like, effective options two, three months later. We saw that happen several times.
Swyx [01:05:54]: By the way, one
Anjney Midha [01:05:55]: Yeah
Swyx [01:05:55]: One thing you also did with the coding agents was that you broke out which are the top coding agents, and they love that. They love that leaderboard. The Klein versus the Rue code versus the what have you.
Anjney Midha [01:06:04]: Yeah. Like, Klein was, like, the top of our leaderboard at the time. We, We then, at the end of. And I’ll skip forward a little bit. The end of twenty-twenty-five, there were quite a few coding apps on the leaderboard, but they were all IDs or, terminal-Agents. And at the end of twenty-five, we saw OpenClaw appear. And OpenClaw was, like, particularly interesting because, one, it was like a new form factor that, like, brought in a new type of user, not just a developer, but like a productivity or a, like an internet creator came to AI for the first time. And it also had an interesting architecture where it was, like, calling your chosen model for these heartbeats to see if it was still alive in addition to using the model for real tasks. And the heartbeats are like, they’re kind
OpenClaw, Hermes, and the Auto Router
Swyx [01:07:02]: Fréquence.
Anjney Midha [01:07:02]: You don’t wanna pay a lot of
Swyx [01:07:03]: Every thirty minutes
Anjney Midha [01:07:04]: To do a heartbeat.
Swyx [01:07:05]: Yeah.
Anjney Midha [01:07:05]: So, the auto router that we provided was really useful to this, like, wide range of users all of a sudden. And so we just saw it rocket exponentially, and then we saw, like OpenClaw just blow up and a couple other, apps lean into that new paradigm and do something similar. Hermes came out and really leaned into things like the auto router and built, like, a really good community and leaned into, like, skill management and making it really easy and effective for people to, like, set their memory in the agent
Swyx [01:07:44]: Yeah.
Anjney Midha [01:07:44]: And build really good skills.
Swyx [01:07:45]: Which another thing you never did, memory skills, sandboxes, all these, like, adjacent things you could have done.
Anjney Midha [01:07:52]: Could have, but It’- I think,
Swyx [01:07:54]: It’s hard to bet.
Anjney Midha [01:07:55]: They’re also - There are things that developer-- that really matter for, like, the developer use cases that were coming out at the time. Like, developers wanted to architect those things.
Swyx [01:08:05]: Right.
Anjney Midha [01:08:05]: Those were kinda critical to building a good user experience. It’s really-- It was, like, - It’s been hard for companies to find abstractions that work for all developers on the memory layer. It is, it - Yeah, there are some, like Mastra has done a pretty good job, for example. But, like, developers have, like, lots of varied preferences for them. And then we - - the way our leaderboard has changed over time is like a movie of how the AI space has changed over time. If you just like, go to the Wayback Machine and look at the rankings leaderboard and the apps leaderboard over time, it shows you, like, what’s happened in AI over the last couple of years.
Swyx [01:08:48]: To me, the coming of age moment was, Andrej Karpathy was like, “I no longer read Local Llama ‘cause, like, I just go to OpenClaw-- OpenRouter’s leaderboard.”
Leaderboards as a Map of the AI Ecosystem
Swyx [01:08:57]: Which I remember that. Yeah. I think he probably, like, said, like, “Sorry, guys, I’m gonna send a bunch of traffic to you.”
Swyx [01:09:03]: So I also wanna bring it into the Stripe, thing.
Why Stripe Acquired OpenRouter
Swyx [01:09:07]: How does that conversation start?
Anjney Midha [01:09:09]: We had this longstanding relationship with Stripe, though, from, like, many different projects that we had worked on with them. We invest, a lot of effort in countering abuse,
Swyx [01:09:24]: Token fraud.
Anjney Midha [01:09:24]: And token fraud.
Swyx [01:09:26]: Can you give some numbers just - so people understand?
Anjney Midha [01:09:29]: I think I, like, I posted about this. We blocked 10x as much dollar volume last month as the month before. And the types of token fraud are diversifying quite a bit. there are, like, fraudsters going after typical stolen credit cards, but there are also, people trying to resell traffic against the terms of service. There’s, like, hacked accounts. There’s people who just lose - like, their whole company is compromised, and they don’t even realize it, and we help them, like, regain control and detect it. There’- There are accounts that are, like, reselling inference on the side. There’- There are accounts that are dealing with, a, like, an accidental runaway agent, and they don’t realize it. Not a hack, but it’s something that blows up and the company doesn’t want it. And so our trust and safety team, like, works a lot on all of these, like, categories of problems and helps block it and detect it. And so we’ve built these. we have models around them. We - We worked closely with Stripe for a while on this, and I think it’s gonna become a huge problem in the ecosystem. Like, we’re already seeing a lot of companies start to see these fraudsters, like, spread and look for other ways other than OpenRouter to other fraud vectors. And if you’re making a gateway or selling, like, generalized inference, you are a target for fraud. If you’re selling very discreet, like, intelligence products that are, like, doing something pretty specific, but not, like, just reselling inference with some added capability, then you’re way less likely to get these fraudsters. So - I think we’ll see companies also move away from just reselling inference with some like, added capability and move towards like, discreet tasks and charging for those tasks and charging for those enhancements and letting people bring their own inference, like, in a party way.
Fraud, Abuse, and the Emerging Token Economy
Swyx [01:11:39]: Whoa. Okay. and yeah, obviously you would power that.
Anjney Midha [01:11:44]: Right.
Swyx [01:11:44]: But you - People pay, for outcomes Or per task?
Anjney Midha [01:11:48]: I think people will pay. I think, like, the Datadog pricing page is a good look at, like, the future to come. It’s like companies, like infrastructure companies will, like, charge for different types of events that they’re providing, and there’ll be lots of, like, continuous pricing models that look like that. And of course, there will be, like, if you go down, towards consumer apps, simpler pricing, more subscriptions, fewer events to worry about, and ones that, like, are not. Focus on just adding a markup on top of inference.
The Token Economy and Security at Scale
Swyx [01:12:28]: Yeah.
Anjney Midha [01:12:28]: Not just because fraud is hard, but also because the pressure from the labs and from - like, good inference providers to, like, do a commit and then bring your inference elsewhere is gonna be very high.
Swyx [01:12:44]: Any comments?
Alex Atallah [01:12:45]: Two. One, I think Alex has done a very eloquent job of describing something, counterintuitively I knew would be a thing at scale, like four years ago because of Discord. And the particular experience that taught me this was, as we started scaling Midjourney, - one of the primary ways that we used to give away or, like, get people to try Midjourney early on to get to their first ten generations. Because, ten generations - ten images generated was roughly the magic moment activation point we found. Like, once you’d done ten, you were like, “This is extraordinary.” but for that week, so we had a free trial with Midjourney. And one day I woke up, because I was the head of platform and had to monitor, I had all these dashboards, and I had, like, three missed calls from David. And it turns out, like, there had been this flood of new users overnight. And we were like, “This is great.” And he was like, “No, we shut down the free trial.” And I was like, “Why is that?” and he said, “I want you to look at the geolocation IP addresses.” And somebody in China had started to resell Midjourney free, subscriptions with the free trial as a way to, like, you - It was fraud abuse, right?
Swyx [01:13:54]: Even for a specialized model like Midjourney.
Alex Atallah [01:13:56]: Yeah. And that was an application. So this idea - I think the big picture realization I had back then was, hey, there’s a new type of unit of value that’s being streamed across the internet called a token.
Alex Atallah [01:14:11]: And over the next ten years, the entire internet value chain was going to have to deal with the fact that, like, the more valuable tokens got, The more bad actors are gonna go to try to get their hands on those tokens. And anytime you scale something and the payload gets more and more valuable, More bad things, people try to get access to that value. And so it was very obvious to me back then. And so, look, to this day, I don’t think there’s a free turn. Like, I don’t think Midjourney’s ever turned on the free trial since then, because it was really not an easy problem to solve in terms of trust and safety. that’s why I - started teaching the class Security at Scale at Stanford. Like, it was like one of - that and the Anthropic learnings, to me, it was clear that the need for security at scale is gonna be enormous a few years from then. Because if you just do the math, right, think about, like, if we’re. online payments, has started roughly in the eighties and nineties, right, and grew to over a trillion dollars over the next ten years, and we needed to build entirely new payment solutions to deal with online fraud. where we are today is roughly there on tokens, but over the next even five years, we’re expecting the token economy to get to, like, roughly 5 trillion dollars. And over the next ten years, I’d be shocked if we weren’t at 10 trillion dollars of token flow. And so if we were starting to see such aggressive abuse and fraud at subscale, Midjourney, remember Midjourney at this point was, like, less than three $100 million revenue run rate a year.
Alex Atallah [01:15:44]: I just realized we were gonna need, like, entirely new, Like, systems to deal with the fraud that was gonna happen for trying to get into the token flow. And so, - I, - I forget the board meeting it was when you brought up that, Stripe wanted to partner up, and it made so much sense to me because Stripe Radar. When I was at Kleiner ten years ago, we invested in Stripe, and the whole pitch that, Patrick and John communicate so eloquently was like, “Hey, unlike traditional payment tools like Braintree that do a day verification, like KYC and AML to get the fraud out of the way, we just bite the fraud cost upfront as customer acquisition cost and - tell a developer, like, just use five lines of code, and we start accepting your payments in five minutes. And what’ll happen is over time, we collect all this data on the developers.”
Swyx [01:16:31]: Cloudflare model.
Alex Atallah [01:16:32]: Is the Cloudflare model, right? And they did. Five years later, they launched Stripe Radar, and Stripe really today is a security company. That’s the real. People think it’s a payments company. No, the reason. There’s lots of other payments providers today that give you, like, cheaper payments transmission. But the reason Stripe keeps, being the dominant one here and Adyen and Europe is because they have extraordinary fraud detection that they’ve built, - over the years.
Swyx [01:16:52]: It’s the same story with Elon and Max Levchin
Alex Atallah [01:16:55]: And affirm, yeah.
Swyx [01:16:56]: Yeah.
Alex Atallah [01:16:57]: So, I think the story shows up over and over again, where every time you have value streamed across the world in large amounts, you need new protection and security infrastructure to fight, to keep the bad guys out and allow the good people to, like, have their transactions happen really fast. And so I think, - this is why - from my perspective, like, the Stripe and OpenRouter story is a security story for the internet ecosystem, for the frontier AI ecosystem. Without a partnership like that, it becomes very hard to defend the quality of experience and the speed and all the good stuff without letting the bad guys get in the way. the second is that, there’s this underappreciated thing about, like, the fact that you need to. Like, - all the bad things that Alex described as being perpetuated by humans right now is going to be perpetuated by AI agents over the next ten years.
Swyx [01:17:46]: Oof.
Alex Atallah [01:17:47]: Right? So think about the, like, recursive scale we’re about to see of bad actors. It’s not just bad human beings, it’s, it’s all the bad agents that are gonna be attacking the token flow. And there’s. It’s very hard if you’re a researcher and at an AI lab to reason about that problem because the only data you have is how agents you’re training are going rogue. But that’s just a fraction of all the bad behavior on the internet that we’re gonna see. And so what you need is defenders, new sheriffs in town, which cowboy hats, that can see all the bad behavior from AI agents across the ecosystem, from different model labs and different trained deployments and different developers, and take all of that data and say, “We’re gonna build a shield for the entire token economy.” Because without that, the amount of fraud we’re gonna see of this 10 trillion dollars in GMV and global GDP growth is, like, a huge percentage of that, I think, is going to be fraud, abuse. And we might never get there if people just don’t trust. Tokens, right? and I don’t think this infrastructure exists. So you have your work cut out for you with, at Stripe, but I don’t think people have realized the scale at which agents, agent, agentic fraud, like bad behavior perpetuated by AI agents is about to hit us like a tsunami.
OpenRouter + Stripe: What Changes Next
Swyx [01:18:58]: Yeah. there’s a lot to dig into there. I wanna give you the last word. We do have to wrap. what can people expect from OpenRouter and Stripe?
Anjney Midha [01:19:07]: I think this is a really good way for us to accelerate market and, to go upmarket more quickly. It’s also, as Ansh eloquently described, this is, there’s a really clear better together story here when it comes to improving trust and safety and making it really easy to, like, accept tokens and let people bring their own inference to your app and to help developers just, like, build on top of inference, going forward. We have a really strong brand with OpenRouter, and we’re keeping the brand. So, like, OpenRouter, like, as a product and the roadmap and the name and the brand, like, is staying the same. And so what, like, you should expect, in the next six months is that most things will be like what we would have done had we been independent, except everything will be moving faster. And that’s like our, term goal. Longer term, hopefully I can comment on it soon, but I can’
Closing: Building the Infrastructure for the Token Economy
Anjney Midha [01:20:11]: Now.
Swyx [01:20:11]: Okay. Well, we’ll hopefully do a follow-up at some point, but thank you for being so generous with your time, and, congrats on the partnership. this is one of the most beautiful bromances I’ve seen in AI.
Alex Atallah [01:20:22]: Just starting out.
Swyx [01:20:23]: Starting from Stanford
Alex Atallah [01:20:24]: Just starting.
Swyx [01:20:24]: To here.
Alex Atallah [01:20:24]: Yeah. Lots more to do.
Anjney Midha [01:20:26]: Yeah.
Alex Atallah [01:20:26]: Lots of sheriff, policing to do of the, of
Swyx [01:20:29]: Yes. The cowboys in town.
Alex Atallah [01:20:30]: Of the token economy. We need We need new sheriffs for sure.
Swyx [01:20:33]: Yeah. Awesome. Thank you.
Anjney Midha [01:20:35]: Thank you.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe- Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time.
One new feature in particular caught our eye: WorldPrompt, a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events. The events, or actions, can even be prompted in real-time.
To understand the implications of WorldPrompt, we spoke to Kamil Sindi, Runway’s CTO, and Robin Kahlow, its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis, co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him.
Who’s building real-time interactive world models?
First, some context about world models that can generate interactive video and audio in real-time.
Runway is reportedly valued at $5.3 billion, based on its most recent fund raise of $315 million in February. Its first release, GWM Worlds, was launched last December.
Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro, and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table:
Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations. For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.”
But as our interviews with Runway show, real progress is being made.
The central idea of WorldPrompt
WorldPrompt, a new feature in GWM Worlds 2, helps differentiate Runway from its competition. You can think of it as a control layer for characters, cameras and the environment. As Kahlow put it, it’s a way to “control all the different subjects in the world” — similar to a computer game.
“Like, if there’s an NPC [Non-Player Character] somewhere, the NPC might walk up to you and say something. So you could achieve the same thing with this kind of model, where you can have very detailed control over everything in the scene.”
As the name suggests, WorldPrompt is a prompting mechanism — not a programming language. So, unlike virtual world games like Minecraft or Roblox, GWM Worlds 2 doesn’t offer scripting capabilities or the ability to control state. But there’s a power to that, as Sindi pointed out.
“You can create promptable worlds on-demand with video and audio in sync, across all these different domains and environments. That’s not a distant-future hypothetical thing,” he said.
But there are also limitations to prompting a world model. We asked how reliably the model would follow an instruction to create, for example, a law of gravity or a certain ability in a character?
“Yeah, so it’s a research preview,” Kahlow replied. “So it’s not perfect, of course, and there are still flaws. It really depends on how difficult the action is. I would say movement works quite reliably.”
Sindi added that more training plus scaling the data and models is resulting in “better following.”
How a video model becomes a real-time runtime
Despite the current limitations of GWM Worlds 2 — especially if you compare it to pre-designed and scriptable worlds like Minecraft or Roblox — the true promise of world models like Runway is that they’ll eventually lead to fully self-generated, real-time games and experiences. Which is an extremely hard engineering problem, as Kahlow reminded us.
“There are two challenges. One is making the model not generate a whole clip at once. So instead, you want it to generate frame by frame while you’re looking at it. And the other challenge is actually making the generation fast, so you can play it in real time.”
GWM Worlds 2 offers real-time interactive worlds streamed in continuous 720p video at 24 frames per second (fps) and audio at 48,000 Hz.
Runway achieved this firstly by taking its foundational audio-video generation model and fine-tuning it to the new WorldPrompt format, so the model can follow that. It then post-trains the model to generate autoregressively.
“And after that, we work on making it real-time through distillation methods,” Kahlow added.
Co-CEO Anastasis Germanidis offered more technical details in our podcast with him. He told us that the process starts from “bidirectional diffusion that basically generates an entire video at once and [makes] it autoregressive.” This allows the model to “generate one frame or a few frames at a time.”
Germanidis described two possible forms of distillation in order to make it real-time: distilling a larger model into a smaller one or reducing its diffusion steps. As a general example, he said a model might go from around 50 denoising steps to four, with some quality loss but potentially comparable results.
The challenges of real-time generation
Germanidis admitted that there were issues with how it generates real-time interactive video.
“The biggest challenge with autoregressive models is error accumulation,” he said. “You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time.”
Sindi told us there are also challenges dealing with “infinite generations” of content.
“There’s all these challenges around what context to keep, what to discard that’s not important. And so there’s all these optimizations we have to think about, so we’re not blowing up our GPU memory.”
Another current limitation is long-term memory. “The model does not have perfect memory,” Kahlow said. “That’s still an open research problem.”
Causality and correctness
While performance is the primary challenge for Runway at this time, its world model also has to produce plausible consequences when a user takes different actions.
Germanidis used the example of simulating football; he pointed out that online video training data contains more successful goals than failed goal attempts, so a video model might render the first more convincingly.
“If I take this action versus this action, you want it to generate equally realistic outcomes,” he told us. “That’s, I think, the big gap between video models and world models: that idea of counterfactual generation.”
Sindi told us that evaluation gets harder the more complex interactions get.
“If you have this multi-prompt, multi-character, multi-scene [environment], how do you really understand what was causal and what was not?”
To try and solve that, Runway has some automated verifiable tests. But since GWM Worlds 2 is a research preview, Kahlow noted that doing tests yourself is also advisable — “trying out your model to see what doesn’t work is really important.”
More than gaming — there are agent use cases too
Gaming is the obvious use case for what Runway is building, but there are others. Kahlow mentioned robotics — for example using a simulated environment to test how a robot works.
Another, more intriguing, use case is to use it to test agents at scale.
“Having thousands of simulated environments is much less challenging if you have a suitable model like GWM Worlds,” Kahlow said.
But how does an agent know what’s changed in the world — is there a structured state that it can read, or is it just the generated video and audio that it’s consuming and understanding?
“So there’s no structured state here,” Kahlow replied. “It’s just observing the same thing you might observe in real life, just [in this case] from cameras.”
Sindi noted that GWM Worlds can also be used for “synthetic data generation for agents.”
Finally, Germanidis suggested there’s potential to use these world models alongside reasoning models.
“You’re maybe using some reasoning [for] planning of the scene, and then you’re passing it into the diffusion head that’s actually generating the pixels.”
Anastasis Germanidis
* LinkedIn: https://www.linkedin.com/in/agermanidis/
* X: https://x.com/agermanidis
Timestamps
00:00:00 Introduction
00:05:17 Runway’s Origins and the Bet on Generative Video
00:12:23 The Stable Diffusion Story
00:18:44 Gen-2, Controllability, and the Weekend Hack
00:23:02 From Video Generation to World Models
00:28:03 Learning From the World, Not Just Language
00:35:04 Sora, Runway’s Existential Crisis, and Gen-3
00:39:39 Why Real-Time Video Is Inevitable
00:43:06 Interface World Models: Software Without Code
00:50:25 The Fully Neural Operating System
00:55:11 World Models for Robotics
01:02:32 Robot Policies and World Action Models
01:07:47 The Lucid Dream Test
01:11:41 Video Agents and Omni Models
01:23:12 Artists, AI, and Creative Workflows
01:27:14 Physical AI and the Future of World Models
Transcript
Introduction: Runway, Creative AI, and the Early Thesis
Swyx [00:00:00]: Okay, we’re here with, Anastassios from Runway, with, me and Vibhu in the studio. Welcome.
Anastasis [00:00:08]: Good to be here.
Swyx [00:00:09]: Congrats on all your success and progress with Runway. You’re opening offices all over the world. Did you envision this when you first started out?
Anastasis [00:00:16]: Not quite. I think even when we started, we had this idea that, It was more a matter of when, not if, we were seeing the early generative models of 2016, 2017, and just extrapolating, assuming, we resolution, quality increases predictably over time. There’s gonna be a point where most of content will be generated, and that was maybe the initial thesis of Runway was we will need, as a result of those generative models, rethink how creative tools are made. and as we built out the research behind, our generative models, it then became clear that they were useful far beyond that as well.
Anastasis’ Background: Art, Simulation, and Machine Learning
Swyx [00:00:57]: And it is more obvious now with, like, the real-world stuff and the world models that we’ll talk about later. I’m just kinda curious how you go from a background in, like, Zocdoc and, computer vision into Runway. Like, take us back to that early conversations with Chris and, whoever else is on your founding team.
Anastasis [00:01:14]: I was always splitting through those two worlds. One was the I had my own art practice. I was making a lot of interactive art, I think for a long time. and then on the other side, I was working in startups, and I was working as a ML engineer, as a backend engineer at different companies. I’ve always been interested in, coding and computation, and especially interested in simulation and brought it back into my early artwork as well. And at the same time, I was interested in
Swyx [00:01:43]: The personal site has a few, right?
Anastasis [00:01:44]: Yeah.
Swyx [00:01:45]: Is there one that we should pull up? Just in case there’s something that’s like. I just like to go down memory lane.
Anastasis [00:01:50]: Yeah.
Swyx [00:01:50]: Okay, what is this?
Anastasis [00:01:51]: So this was, a project that I made, I think back in 2015, where I built this software that would give, voice instructions to people in a gallery space. So it would coordinate interactions between people. And so it will first give you an identity, like you’re an, architect, you’re 30 years old, and, you like sports. and then it would match you with another person, and you have this completely generated interaction. language models were not quite there at the time, and so it was it was a mix of some templates and some, like, some Markov chain-generated text, and it would just completely simulate these small talk conversations between, everyone in the gallery space. so was always very fascinated on the one hand with, generative models and, like, the early machine learning work that was being at that time. But at the same time, there was this separate thread of simulation and what it means. Like, what can we learn about humans by creating those very simple models of their interactions and their behavior?
Early Generative Art: pix2pix, GANs, and Uncanny Valley
Vibhu [00:02:56]: Did you generate the prompts or, the 30-year-old, whatever? Was it you generating them? How’d you, how’d you
Anastasis [00:03:03]: Exactly. So the program would just generate- those, from. Yeah, a lot of it would be Mad Libs style of just
Vibhu [00:03:10]: Yes
Anastasis [00:03:10]: You have lists of different professions, lists of different,
Vibhu [00:03:14]: Hobbies
Anastasis [00:03:15]: Personality types, lists of different, ages, things like that. And then it would just combine those things together. And then maybe the next project we go is, Uncanny Valley, Uncanny Road, which was
Swyx [00:03:27]: Gans
Anastasis [00:03:27]: One of the first projects that, we built with, one of my two co-founders, Chris. This was taking, pix2pixHD, which was one of the early image-to-image models that NVIDIA released back in 2016 or 2017. and it was a model that would take a semantic map of a scene and then generate a photorealistic, let’s call it, output. very early days, so it was not very high-fidelity outputs, but it w I think was the first image-generation model that could generate at 1K resolution. And it was all trained on self-driving datasets. So the semantic categories it would support were only, things you would encounter on the road. So it would be pedestrians, traffic signs,
Vibhu [00:04:16]: Stoplights
Anastasis [00:04:17]: Bikes, stoplights. And so that was one of our first indications that we built this and people were making all this, like, very surreal imagery of, yeah, a million plus a million pedestrians or a million traffic signs or, like, gigantic humans. And it was a indication that you could take a model that was trained on this very boring dataset, essentially, of, like, not that many interesting things happen when you’re on the road, and then you can repurpose it and go very out of distribution and make something that was artistically compelling. And that was It’s a summary of the thesis of Runway in some ways, that you can take the same generative models, and if you look at them from another direction, if you build interesting tools around them and you give them to artists, they’re gonna do things that you don’t expect.
Vibhu [00:05:02]: Very cool. I like the, UX of it. You’re just given an empty canvas, try whatever, do whatever. And then the other one, like, you see everyone with wired headphones? Like, that’s, that’s a sign that it’s, it’s very
Anastasis [00:05:16]: The Apple
Vibhu [00:05:17]: Yeah
Anastasis [00:05:17]: Apple, your version.
Vibhu [00:05:17]: Original ads. Yeah. Take us to today. You’ve been doing this for seven years at Runway. How have we got to this? Like, how do we go from driving simulator data to all this? And you cover the whole stack of generative media?
From Creative Tools to a Research Lab
Anastasis [00:05:33]: Interestingly, we’re almost back in, we’re, we’re full circle. We’re, we’re now applying our models and beyond creative tools into real-world scenarios. But it was a, it was a long journey. It was very early on we realized the first version of Runway was a way to easily use the, all the open source model of the day, things like pix2pix to. and give them to artists. That was the initial idea, is those models are too difficult to use if you’re not a machine learning engineer. Like, what happens when you give them to artists? Very quickly, we realized we needed to build a research org, inside of Runway, and that happened maybe on year one. And, a lot of the mandate there was. The image-generation models of the time, the video generation models of the time, or there were barely any video generations all the time, but they were not quite there where they could be productionized and brought into tools that would be part of creative workflows. so we need to push the frontier of the research. And so maybe the first four years of Runway, research was almost happening on the background until there was a moment in 2022, with latent diffusion, with, DALL-E 2, where, there was that step function change, and you guys maybe remember around the time.
Swyx [00:06:49]: I started in this space because of latent diffusion and Stable Diffusion.
Anastasis [00:06:54]: Yeah.
Swyx [00:06:54]: Because I was like, “Wow, this is not only, like, feasible, it is doable on consumer hardware.”
Anastasis [00:07:01]: Exactly, yeah.
Vibhu [00:07:01]: I think the delta is also huge. Like, I learned pix2pix. Like, this was intro to ML, the TensorFlow, like, Jupyter, Google Colab notebooks were like this, and then you have a sudden step function change, with diffusion and whatnot. Any other ones since that. Like, there were clear examples of what early diffusion were to get to here. Any other changes in key technology research?
Green Screen, Rotoscoping, and Early Runway
Anastasis [00:07:26]: Between, 2018 when we started and 2022?
Vibhu [00:07:29]: Yeah.
Anastasis [00:07:29]: So one of the early work that we did in Runway was solving segmentation, image and video segmentation. It was a very important problem because most VFX involves essentially separating
Swyx [00:07:42]: Rotoscope
Anastasis [00:07:42]: Subjects. Yeah, rotoscoping. Extremely manual process. Nobody enjoys doing that. and so a lot of the early days of Runway was building this tool. It was called Green Screen, and it was for a long time the main thing that people were using Runway for. It ended up being used in, Everything Everywhere All at Once and a bunch of other high-visibility films and series. But that was essentially, Runway for a long time was a post-production tool until latent diffusion and generat- Gen-1, Gen-2, happened.
Swyx [00:08:12]: Cool. let’s, let’s go past that moment. You’ve come a long way. Then you started releasing your own models. Maybe describe that journey as well.
Scaling Video Models and the Bet on 1,000 A100s
Anastasis [00:08:20]: Yeah, so we go to the other point, yeah, in mid-2022 when it became clear that we’re doing research at a fairly small scale of compute, and it became clear that, like, scaling laws would apply to, image and video gen in the same way that we’re applying to language generation. So we made a big bet, and I think at so at the time, we signed this deal to build a cluster of a thousand A100s, which at the time we were a Series B startup. That was a almost, slightly irrational decision maybe, but we really believed that if we trained a video model at a large scale, we would get, like, a great model at the end. And at the time, the goal or we set the goal around fall of 2022 of what is, what does the latent diffusion, Stable Diffusion moment look like for video? And at the time, the best model of the time was called CogVideo. it was one of the early video models. It was very 256 by 256 resolution, very not very high quality. and so we decided we’re gonna build out this cluster, and we’re gonna just invest in, like, in building out our own video model. it became clear as we’re training Gen-1 that it was difficult to get to fully. we wanted to build text-to-video, but it became clear to us that an easier starting point would be to start from video to video. Because when you have a stronger conditioning, it’s, it’s an easier problem to restylize an existing video versus generate the video from scratch. And so we released Gen-1 first back in, it was January of, 2023. Yeah.
Vibhu [00:10:04]: It’s just a fun visual podcast, honestly. Like, if we can see February 2023, what was the state of stuff?
Gen-1: Video-to-Video and Depth Conditioning
Anastasis [00:10:10]: It’s so interesting ‘cause at the time when you see those results, you think this is so incredible, and this is like, it’s almost like image generation or video generation is solved. And then you look back a few years after, and it’s like, it’s It’s just like you get used to the results very quickly, with those models. But at the time when we started seeing those results, it was, it felt quite incredible, and the level of, like, quality that you could get. And, so the Gen-1 was a depth-conditioned video model, so it would turn. it would take a input video, it would predict. it would it would first convert it into the depth map, and then we would generate, pixels with a latent diffusion model.
Swyx [00:11:01]: Yeah, very effective.
Vibhu [00:11:02]: Yeah. I didn’t realize how distracting the blog post would be. Sorry.
Anastasis [00:11:05]: Yeah, but, one of my favorite examples of on those, on Gen-1 was both, if you go up to mode three or mode two, there was this storyboard use case where people would make
Vibhu [00:11:18]: Ooh
Anastasis [00:11:18]: Would
Vibhu [00:11:20]: You can mess around with the
Anastasis [00:11:20]: Make a city out of books or out of boxes, and then they would shoot a video with their phone and then translate it into a photo-photorealistic output. There was all these ways in which those models were starting to be used for storyboarding and also for really. and then if you go to mode four, like, of taking untextured 3D scenes and then turning them into photorealistic output. So we saw a lot of use cases early on where people that were familiar, were power VFX editors would just take a blender, render, and then they would get translated in with Gen-1 or create a scene in Unity and then take a capture a video of it and then translate into, restylize it. So I still think video to video is powerful. I think we had a recent video-to-video model as well, and it’s one of my favorite ways of using those models is essentially using them to use ground truth video as, like, the initial inspiration and then translate into different styles or different outputs.
Stable Diffusion, Stability AI, and Open Source
Swyx [00:12:23]: But I think we’re gonna go into, like, the rest of Runway and catch people up to speed today. I did wanna cover the, let’s call it the Stable Diffusion controversy, or, what happened with Stability AI, whatever. I think there was a two sides of the story. I think there’s part of that is a normal thing of, like, people, join and leave companies, but what is the, retrospective now that, there’s been some years behind it?
Anastasis [00:12:49]: Yeah, it’s a very, it’s a very long story to go into. I think it would
Swyx [00:12:53]: Which I remember you wrote a really long post about.
Anastasis [00:12:56]: We would probably cover the whole hour to go into it in more detail. But, essentially, there was the latent diffusion paper that came in, I think that was at the end of, 2021. And then Patrick Esser, who was one of the researchers behind, latent diffusion, and he worked at Runway at the time, he built latent diffusion in collaboration with Robin Rumbach and a few other folks back, in the in, CompVis, which was, a lab
Swyx [00:13:26]: Like a research group, yeah.
Anastasis [00:13:27]: And, after releasing the early latent diffusion model, they, essentially they were. the goal was to keep working on versions of the model, scale it up, incorporate new data, incorporate new tasks. And Stable Diffusion was the same model, but trained on more compute, and then with a few more tricks, like a classifier-free guidance paper came at some point, I think in the early 2022. And that
Swyx [00:13:52]: Which, like, was a big prompting improvement.
Anastasis [00:13:55]: Yeah.
Swyx [00:13:55]:?
Anastasis [00:13:56]: That improved results. it was trained on better data, so like, the esthetic subset of LAION, but it was effectively, the same underlying architecture. And there was that big training run, that, happened on Stability’s cluster. Stability financed that run. And looking back at that story, I think it was the work to build and train that model was done. It was a, it was a research project. It was done as part of, like, continuation of the latent diffusion work. It then, I think it the model became very successful, and it, I think there were the. And I think as a result of its success, other companies tried to, figure out the commercialization path for it. But for us, it was very important that we try to, we make sure that we. It was meant to be an open source research project, and so the we decided that we should continue releasing versions of it, since that was the original goal of Stable Diffusion, and that led to releasing Stable Diffusion 1.5. There was maybe a day of, a bit of, miscommunication there, but ultimately that was resolved very quickly within hours. so yeah, there was
Swyx [00:15:12]: Okay
Anastasis [00:15:12]: Not a nice
Swyx [00:15:13]: I just wanted to. you have to
Anastasis [00:15:15]: Yeah.
Swyx [00:15:15]: You’re one of the main players in that journey, and so it’s nice to hear from the source of, like, what happened. Yeah.
Anastasis [00:15:22]: Yeah. I think it’s all, it’s all in the past now
Swyx [00:15:26]: Yeah
Anastasis [00:15:26]: I would say. and, like, both companies, Stability took its own path, Runway took its own path.
Swyx [00:15:32]: Yeah. There’s still. James Cameron is backing the new Stability, whatever they’re doing with the Hollywood studios.
Anastasis [00:15:38]: Right.
Swyx [00:15:38]: I don’t know what they are doing. I think one thing that impresses me, and I’m happy to move on, is that back in the that time, let’s say, like 2021, 2022, there was this community of people that you were involved in that was researching all this stuff, right? And, like, from everyone I talked to who was active then, it seemed like it was fairly obvious that somebody would do the hero training run that would produce Stable Diffusion. So, like, I guess the question is, like, you had the you were you had made investments. You were you had the foresight. Is it accurate to say, like, that is reflective of, like, what people were thinking at the time? Or was it still very much like, “Well, we’ll use it as, like, a post-production tool or something. I don’t know.”? Like, where in the sentiment were we that maybe you can think back to, like, what the community was like back then?
The Early Creative AI Community
Anastasis [00:16:28]: I reminisce and I think very fondly those early years, from like 2018 to 2022, because it was a very small community that, as you said, were very convinced that this was gonna be a big thing. And at the time, anyone who. Because it was such a small circle and, everyone who would, like, be part of that circle and, like, make projects with it would, immediately get, go viral. so like
Swyx [00:16:55]: And you didn’t know who they are, right? They’re just some name on a, GitHub or Hugging Face somewhere.
Anastasis [00:16:59]: Exactly, yeah. So I remember one of the first big viral moments of creative AI was, there was the neural style transfer paper
Swyx [00:17:09]: Huh
Anastasis [00:17:09]: That
Swyx [00:17:10]: Something dreaming?
Anastasis [00:17:11]: I think it was called neural style transfer.
Swyx [00:17:14]: Okay.
Anastasis [00:17:14]: There was also Deep Dream, the puppy slice
Swyx [00:17:16]: Yes
Anastasis [00:17:16]: Which was, also really cool. but, yeah, there was this project that, Jim Kogan, who was an early advisor of Runway and one of those,
Swyx [00:17:25]: Marketing guys
Anastasis [00:17:26]: Big, creative AI, folks, he literally just, like, showed a video of himself taking the New York Subway and going over the Williamsburg Bridge and then stylized it with, I think in the style of Van Gogh or, like, one, painter. And that was. Like, at the time, that was, like, so cool and it went viral and it was completely revelation to people that you could do this with generative models. And that was only, it was less than. It was maybe 10 years ago. So just, like, as an indication of, like, how quickly things have gone.
Vibhu [00:18:02]: It’s pretty crazy. Like, even since then, you’ve got people at every level of the stack. You’ve got devs, creatives, artists, hobbyists. You’ve got everyone using it. And for people that tried stuff early, they’ll remember how hard it was to use regular diffusion, right? Like, nowadays, you can use your favorite ChatGPT image gen or whatever, give a sentence, get a beautiful output. But diffusion was like, the whole ultra HD, 4K, high resolution. Like, prompting these things was very different. anything you learned on the tooling side, like from the offerings you guys have now, so like creatives, devs, you really took the. Research and brought it to everyone to use. anything interesting there to share?
From Gen-2 to Controllable Video Generation
Anastasis [00:18:44]: We had to build the entire model serving infrastructure for video diffusion models. There was nothing else, already, like, because we had Gen-2 was the first text-to-video model, I think, out in the market. So many things that we learn over time. I think the I think the biggest one was, like, we. it was very clear early on that text-to-video was not gonna be the answer. Like, you. Like, people wanted a lot more control than that, and so we invested in, like, control building on top of those models very quickly. how do you use the camera trajectory as control? How do you use an initial input frame as control? So that was a very early learning for us. With text-to-video was, like Gen-2 was an amazing, step function improvement in the quality of video models, but it was used much more in an exploratory way because there was nothing to ground it to. There was no reference that you could bring into it. There was no. You couldn’t really control the camera motion. You couldn’t control the object motion. And so the first year, in 2023, was really all about what are all the interesting ways in which we can condition those models? And it was a lot of just post-training rounds on top of the base model to figure out, like, what, -- how do people wanna control them? And so there was, like, this quick succession of the we it was called Motion Brush, which was you could, like, you could draw arrows and dictate where things should move in the scene.
Vibhu [00:20:09]: That’s so cool.
Anastasis [00:20:09]: There was camera control that was you could just describe, like, how you want the camera to move in the scene. And because we work with filmmakers from the most of the history of Runway, we immediately got this feedback and got this, decided that this was worth investing in. And so control ability became a big theme, I think, very early on as we were building, as we were building those models. Something fun that I haven’t really talked about too much was just how Gen-2 came to be out of Gen-1. So it was a bit strange because we announced Gen-2 two months after Gen-1 and
How Gen-2 Came From a Weekend Hack
Vibhu [00:20:43]: We’re accelerating.
Anastasis [00:20:44]: It was before Gen-1 was even generally available. But Gen-1 was a depth-to-video model, so it would take a depth map and it would convert it into RGB. and we couldn’t get, text or image-to-video to work directly, and that’s why we started from depth to video. but, and we had discussions of like, okay, we need to spend the next six months investing in text-to-video, maybe increasing the compute scale or the model scale, like train a larger model. And I had this weekend project idea, which was, what if I take a model that, starts from text input and converts to depth maps and then use Gen-1 to convert the depth maps Into RGB?
Vibhu [00:21:29]: It would probably work.
Anastasis [00:21:30]: And so Gen-2 was that.
Vibhu [00:21:32]: Oh. The hackathon pipeline.
Swyx [00:21:35]: The weekend hackathon pipeline.
Anastasis [00:21:36]: Yeah.
Vibhu [00:21:37]: But it looks good.
Anastasis [00:21:38]: And it worked pretty well. there were if you, with the knowledge that it has this, like, two-stage pipeline, you can tell in some cases that the structure of the video looks a bit off because you had to generate the depth first before you go into the output video. But it worked and it allowed us to bring this to our, to users very quickly. But it’s now it’s interesting because, like, people are coming back to this almost two-stage approach. Like, if you look at the Reve text-to-image model that came a few months ago, it had this planner model that would generate bounding boxes before it fed that into the diffusion transformer.
Swyx [00:22:19]: Yeah, Ideogram also the same day.
Anastasis [00:22:22]: Yeah.
Swyx [00:22:22]: I remember that was very strange that both of them came out the same day with the same exact innovation.
Anastasis [00:22:26]: It’s a small community, I think.
Swyx [00:22:28]: I’m like, this is like, this is completely coincidental, right?
Anastasis [00:22:32]: People talk. So yeah, there’s, there’s definitely something into this approach. And, now, like every single like, video generation model in production uses a complex prompt completion pipeline under the hood. I think that’s no secret that there is. That
Swyx [00:22:48]: Humans are terrible at prompting.
Prompt Rewriting, Camera Control, and the Seed of World Models
Vibhu [00:22:51]: I think across the board.
Anastasis [00:22:51]: Yes.
Vibhu [00:22:52]: But yeah, I think like the original Sora one blog post even told you that what happens after your input is rewriting your prompt. It’s much more descriptive about what you would want.
Anastasis [00:23:02]: Exactly. I, And there was the DALL-E 3 paper beforehand that, was the first public, description of the fact that synthetic captions and really detailed captions work really well. And then Sora built on that. Yeah, so it was 2023. We were releasing all these updates to Gen-2, like the camera control, Motion Brush. And there was something very interesting about camera control because it was the first time that you felt that instead of, like, you were creating video, you were creating a short video, you were navigating inside the world. And I think camera control was maybe the seed of some of the ideas that we had around world models and really opening up that research direction. We realized, it was this era and this series of, Gen-1 and Gen-2 models really proved to ourselves, yeah, this is the
Swyx [00:23:56]: Cool.
Anastasis [00:23:57]: So this is not the original camera control. This was the updated camera control on top of Gen-3. But yeah, I think it made those models usable to filmmakers, I would say. The so camera control was very popular. And so we realized, there is one way of seeing those models, which is, you’re just as content creation machines, and there is the other way, which is you’re. As you’re predicting video in order to predict video well, you need to simulate the world in an increasing and increasing capacity. And if scaling laws apply on video, just like they apply on language models, then as we scale the compute that we put into those models, then they’re gonna be able to simulate physics, they’re gonna be able to simulate human actions and dynamics increasingly well and predictably well. That was the thesis about around our efforts on world models, and we spin up this research group to just focus on the world models and how do we turn the video generation models that we’re building into something broader and something that would be useful beyond, also content creation as well.
Swyx [00:25:04]: And that was roughly when?
Anastasis [00:25:06]: Yeah, so that was in
Swyx [00:25:06]: Oh
Anastasis [00:25:07]: In late 2023.
Vibhu [00:25:08]: Interesting. like, I think, a lot of people have been saying a lot of video gen model companies have all pivoted to world models these days, but like, 2023, you’re posting it. one
World Models: From Video Generation to Simulation
Swyx [00:25:21]: It’s, it’s debatable whether it’s a pivot.
Vibhu [00:25:23]: Yeah.
Swyx [00:25:23]: Like, arguably
Vibhu [00:25:24]: Yeah
Swyx [00:25:24]: That’s what you always had to do anyway, right?
Anastasis [00:25:26]: It’s in a way an expansion
Vibhu [00:25:28]: Yeah
Anastasis [00:25:28]: Of the applications
Vibhu [00:25:29]: Yeah
Anastasis [00:25:29]: Of the models as they become more capable.
Vibhu [00:25:31]: The early signs, it seems like the original models you guy had, guys had, people would say it’s very not bitter lesson pilled, right? You’re adding, rewriting prompts, you’re having all these one-off things, but that’s just the state of the tech as it was versus the future of as you said, you can scale it up as, we can scale up to world models.
Anastasis [00:25:50]: Yeah. So it just became. And if you looked at the outputs of Gen-2
Vibhu [00:25:56]: Yeah
Anastasis [00:25:56]: It was not. I think it was not obvious to people that this would scale to become a general simulator of the world. Like, you had very limited movement, you had, very low fidelity or low resolution, like obvious mistakes in human anatomy, like all kinds of limitations. But it was just, the idea was that’s just GPT-two, and GPT-two, it can barely generate, like, coherent sentences. Similar, Gen-2 can barely create coherent video, but if you scale it up, you’re gonna. There is no reason why it shouldn’t work in a way. It’s, And I think that was. That’s, that’s always the mindset of Runway is like this extrapolation of, like, if, like, even when we started in 2018 and you looked at the results of the day, you need to look more at the trend of, like, where we were in 2018 versus when we were at the, when the first GAN came out in twenty, four 2014 or twenty, fifteen. And, you started from, like, thirty-two by thirty-two images of faces, and then by the time in 2018, you could generate, street images at the 1K resolution. And it was the same with world models, very early signs of something much bigger.
Swyx [00:27:08]: Yeah. I was gonna say, like, it’s diffusing into focus. Like, if you look at our visible output from year to year, it looks like a diffusion process itself.
Anastasis [00:27:17]: Yeah.
Vibhu [00:27:17]: Especially watching the early, like, old blog posts, you can really see the choppiness, the details.
Anastasis [00:27:24]: Yeah. Like human civilization starting from random noise and then
Vibhu [00:27:27]: Yeah
Anastasis [00:27:27]: Denoising into
Swyx [00:27:28]: Yeah. Just run it a hundred years.
Anastasis [00:27:30]: Civilization.
Swyx [00:27:30]: Yeah.
Vibhu [00:27:31]: That’s how you’re on track, you’re still noising, right?
Swyx [00:27:34]: Yeah. I like the way that you guys phrased it when you, announced it in June, which is, oh, that you had a video essay. “The human mind is no longer the center of AI. Our world is.” Right? Which is, let’s, let’s call it the past five years of LLM-based AI is very much like trying to emulate human preferences and human speech. But now that’s, like, mostly solved. I think that’s, like, some of the context of your essay, which you also wrote around the time. And now it’s like the focus is on modeling the world accurately.
Scaling Laws for Video and Why Predicting Pixels Matters
Anastasis [00:28:03]: Exactly, yeah. So the way we see it is, there is that, initial mission statement of DeepMind, which is, solve intelligence and then use it to solve everything else. But I think it’s starting from everything else, could be valuable of, like, starting from. there is just so much complexity, and detail in the world that in order to. That it’s, it’s hard to learn directly from just human descriptions of the world. Like, we’re assuming that, like, language models learn from everything that humans have written about the world, like our own understanding as of, the twenty twenties. And there is just so much that we don’t know and so much that’s not captured by existing text, about both the low level dynamics of the world, like we’re not describing in detail. if I tell you to describe, like, how do you tie your shoes, that’s a very difficult thing to describe in words, but it’s very obvious thing to demonstrate. And so I think there’s been. And there’s, more of X paradox, like we’re constantly underestimating all the complexity that goes into very, like, things that we do subconsciously as humans, and we don’t even necessarily always have the words to describe them. And so in my mind, the simulating the world and simulating, physics, simulating the dynamics of the world has always been underestimated, compared to, we place too much emphasis on the things that are easy to talk about. but there is just all this complexity and richness of the world that if we just try and train directly on that observational data instead of training on how people describe the world, we would learn something new that we wouldn’t otherwise know.
Swyx [00:29:54]: You think that the present architectural paradigm is fine? You don’t need, like, another layer, like JEPA, like another famous, New York AI leader would say?
Anastasis [00:30:05]: We’re a very pragmatic research lab. If, we have evidence that an approach works better than the approach that we’re taking, then we have no qualms to taking it. We just have seen no indication that video prediction itself doesn’t scale. And even if you look now, not just our work, but the work of others, you’re seeing in robotics some of the most promising work, starts from video prediction models, and then you adapt them to also the action models, for example. so there is very little evidence that you need something else and that your time is better spent on a novel architectural change compared to improving data and improving the, and scaling the current approach. And so, We don’t have any indication that. the, there is that counterargument that I think there was a tweet by Yann LeCun a few days ago that, understanding the dynamics of the world is very different than, generating, cute videos.
Swyx [00:31:05]: And your answer is no, they’re the same thing.
Anastasis [00:31:07]: Yeah, they’re the same thing.
Swyx [00:31:08]: My cat videos are the same as understanding physics.
Anastasis [00:31:11]: Right, because if you wanna generate. video models can cheat and, like, they could you could give, like, successive dif shots of the scene in a way that doesn’t require you to simulate difficult physics. There is like, all these different ways in which you can hide the deficiencies of the model, and it’s important not to be too tricked by the performance of the current video models. It’s easy to, cherry-pick examples and think that video models are further advanced than they are. So there is a lot more work that we need to do to improve those models. But in my mind, very similar to language, and, like, we’ve. you go from barely coherent sentences to something that, could hold a conversation with a human to something that could can operate autonomously for a day and, like, create entire code bases. And the main difference, there is some architecture improvements along the way, but the main thing is scale. And so it’s the same bet for video, and we have no indications that this is saturating. Like, we have benchmarks that we use for measuring the physics of those models, and we see those predictably improve as we scale those models. So there is. If you want to Google up, Physics-IQ, is one of those benchmarks that measures how well does the model perform at solid mechanics or fluid dynamics or optics.
Vibhu [00:32:32]: I’m curious if you’ve seen any emergence, any scaling law around this.
Swyx [00:32:37]: Yeah, he’s saying there is a scaling law, right?
Anastasis [00:32:39]: Exactly.
Vibhu [00:32:40]: Yeah,
Anastasis [00:32:40]: So the way those models, those benchmarks work is you. the researchers have gone and, like, captured, a few videos that are representative of different physical phenomena, and then you can take the first frame and then pass it through an image-to-video model and then generate a rollout that shows what should happen next. So you have, a ball hanging from the ceiling, and then you use that as input, and then you the model predicts how the ball should fall on the ground. and this measures. we have an intuitive understanding of physics. I know, you can imagine what will happen next if I drop this bottle. So it’s measuring that same intuitive physics understanding of those models, and we’ve measured that at different model scales, and we see, and compute scales, and we see that the score on physics IQ predictably improves. There’s other, tricks and techniques that you can make to improve the score even further, but even scale alone helps, in the model learning better physics.
Swyx [00:33:40]: My main sympathy with Yann LeCun is the, Plato’s cave allegory, right? Like, you’re, you’re, like, learning on the output of a thing, not the internal process of a thing, and it’s very noisy. And, if only you could observe the internals of a thing. It’s hard to observe the internals of a human mind, but you can very much observe, or at least we have a whole branch of science and physics that we’re ignoring on how to model Physics and movement and, gravity and, other interactions. and we’re just, like, throwing away all of that and just saying just scale data, which is very much the lesson of unsupervised learning, but it feels wrong. that’s the main idea.
Anastasis [00:34:21]: I think the history of machine learning is, at large, it feels wrong.
Swyx [00:34:25]: Yeah. It’s a bitter lesson, right? Yeah. It’s, it’s, it’s the simple answer to that.
Vibhu [00:34:29]: I guess, how much can you scale? So, like, even on, let’s say, the video generation side, like, there’s one side of video understanding. Video generation, are we still gonna have tools where it’s like, I wanna generate two hours, twenty hours? there’s a infra way to do it in batches and stitch it together, but, like, do we just keep scaling? Do we just continue long generation consistency, all that at scale? And, like, tying it into where we’re at now from we looked at Runway two to four point five
Gen-3, Sora, and Runway’s Scaling Inflection
Anastasis [00:34:58]: Yeah.
Vibhu [00:34:58]: Like, technically, what advancements have we made to today, and then where do you see things still going?
Anastasis [00:35:04]: So part of the answer is definitely scale. and that was. We learned that lesson in a big way for with Gen-3. So Gen-3 was the model we released the year after, like in 2024. That was a few months after Sora was released. so yeah, there’s an interesting story of that came to be as well. Gen-3 for us was, the first time that we really needed to build. we had to learn all the lessons that the language model world learned in two in three years in the span of a few months. one of the biggest changes of Sora was using diffusion transformers instead of convnets. So a lot of the early, latent diffusion models were all, convnets for the diffusion model part. And the diffusion transformer paper came at some point in 2023, and it showed scaling laws for image, diffusion transformers. And we realized at that point that we needed to invest in infrastructure for model parallelism, for really scaling training to larger than, a few billion parameter models. And we spent maybe the, most of the fall of 2023 building out our infrastructure for distributed training. And we had a lot of false starts and a lot of failure in trying to scale, image and video diffusion transformers. And at that point, February 2024, Sora comes out, and the results are
Anastasis [00:36:35]: Very much superior to what Gen-2 could produce. There were a lot of, a lot of chatter on Twitter about Runway. Runway’s done. like, there is no way Runway will catch up. And if you remember, also OpenAI in the early twenty-It felt very, like it’s a
Swyx [00:36:56]: To the moon
Anastasis [00:36:57]: It’s a formidable opponent now, but at that point, it, they were on the top of their game. nobody could even get close to them. There was maybe Gemini was just the first version of Gemini had just released. So when OpenAI came with Sora and it was such a big jump of like quality, it gave me, there was like an existential crisis for a few hours. But that, I think the amazing thing about Runway and like I think the, we’ve been around eight years now, which is almost we’re dinosaur in AI, and we had to like, we had there was a lot of those moments we had to learn, adapt very quickly and build out skill set in the team that we didn’t have. And so, if you ask anyone what is their favorite time at Runway that was there during that time, it was that push in like three months to get to a model better than Sora. and it, we scaled 10x the model scale, the model size and the, compute that we were training on. we figured out model parallelism. We had zero expertise in that. And then we came out with Gen-3 during that summer. So that was a big turning point, I think, for the company where the research org grew very quickly, and we really started pursuing this vision of the general world model, in earnest, I think after Gen-3 was out.
Swyx [00:38:12]: Yeah. that’s the amazing thing about building when you’re building. There’s no stack to. You have to invent everything yourself. You have to be completely full stack. Now I think like there are inference specialists like Fal or whatever that can help with like, model serving, and I think you guys work with them as well. but yeah, like it’s, it. But at the time, it was just. It’s very interesting to think about what you do when Sora comes out and people are questioning whether your company should still exist.
Distillation, Turbo Models, and Real-Time Video
Anastasis [00:38:41]: Yeah. And yeah, there was no, there was no VLM of diffusion models. Like, we had to build the whole model serving infrastructure and make things efficient. And a few months after we released Gen-3, we released the Turbo version, which I think was the first step-distilled model in production.
Swyx [00:38:56]: That was a whole trend that we covered as well. Yeah.
Anastasis [00:38:59]: So that allowed us, to serve those models at the larger scale, ‘cause I think the first version of Gen-3 was quite, expensive to serve.
Swyx [00:39:09]: I think the whole like trend in like consistency models, Lightning and, Turbo and all these things somehow didn’t really stick around. I don’t know if you have any reflections on this. Because at the time, I was like, “Well, everything should start with a distilled model first, and then you can upscale,” right? It. your bigger models just turn into fancy upscalers, but like you should always draft with a smaller model and faster model, right? Because you can get it so quickly, like near real-time.
Anastasis [00:39:39]: Yeah. I would not be so sure to say that didn’t stick around. I think that, it’s, it’s likely to. that there is a lot of step-distilled models that are actively used in production. there is still a gap in quality compared to the, non-distilled model. but in my mind, we’re still. there is a two to three year offset from language models. So the things that, So it’s just a matter of time before there is better distillation techniques. we use. Right now we have a real-time model core character that I think is the largest deployment of real-time video models, that’s a step-distilled model, and it’s actively being used. It’s a very specific use case compared to a general video model. So this is a
Swyx [00:40:27]: Very cool, by the way.
Anastasis [00:40:27]: This is avatars stuff, right?
Swyx [00:40:28]: Consistency, character.
Anastasis [00:40:30]: Yeah. So this is a talking avatar, model. we were able to. we optimized the hell out of it, and it generates at 24 FPS, and it’s a, it’s a step-distilled autoregressive video model. So if we look at our world model direction, a big component of it is starting from the bidirectional diffusion that generates entire video at once and making autoregressive shows. So you generate one frame or a few frames at a time. so there’s a lot that goes into that pipeline of getting to a real-time model. It’s first you need to make it into a causal autoregressive model, and then you just turn it into. You need to do some additional step distillation to get it to be real-time. and I think that part is just starting. I’ll be very surprised if we’re, two years from now, we don’t primarily use real-time models. To me, real-time video generation is just inevitable that, it has much better user experience, it’s much cheaper to serve, and, the quality gap between the base model and the real-time model is only gonna close as we figure out better, distillation techniques. And we made a lot of progress there internally on maintaining the quality of the base model when we distill them.
Swyx [00:41:49]: How much of this is transferable? So is it the same base model? Like if you’re doing diffusion across the whole sequence and you’re converting it to step autoregressive distillation, is this like distillation where you still need to train both, you can use the same base and converter? What’s that process like to go from regular model to something that’s real-time on a technical level?
Anastasis [00:42:11]: So the nice thing about diffusion models is you have, two axes of distillation. So there is the. You can distill to a smaller model, which resembles what you do in LLMs, or you can distill in terms of taking less steps, less diffusion steps. So you could take a model that generates in fifty steps and generate in four steps and get to, You have some performance, degradation, but very often you get comparable outputs. So you can even take the large frontier model and distill it with step distillation and get to a real-time performance, and that’s what we’ve seen. So, depending on the use case, in some cases we might also serve with a smaller model, but in a lot of use cases, we just use the
Swyx [00:42:56]: Step distillation
Anastasis [00:42:56]: The frontier model, and we’re able to make it work in real-time.
Swyx [00:42:59]: I think this might be a good time to cut over to his laptop to show off some of the real-time stuff that you’re doing.
Interface World Models and Neural Software
Anastasis [00:43:06]: This is one of the research updates that we did recently. so we’ve been working and f in getting our general world models to, different applications. one of them that we think is very compelling is using general world models as essentially, an interface, a universal interface to software. This is a version of our world model that’s called an interface world model. and the idea is that it essentially, replaces, the, front end of a software application. It renders the pixels directly of an interface and is trained to predict what happens next as a result of, a click or another interaction you have with the interface. So this is all pixels. it’s there is no HTML, CSS, React that’s powering this interface. This is directly at the output of our real-time, video generation model, and it takes clicks directly as input.
Swyx [00:44:09]: And drags, click and drag.
Anastasis [00:44:12]: Right. So it supports
Swyx [00:44:13]: Ooh.
Anastasis [00:44:14]: Yeah, clicks. It supports drags. it also supports scrolling. and the amazing thing about this is that you can effectively describe in the prompt how you want different elements, like what do you want the behavior of different elements to be. So it’s almost you’re you can turn, an interface from, markup language description of, like, an HTML interface, and instead you can just describe the interface. if I press this button, I expect this to happen. If I press this button, this should happen. And it’s useful, we believe, both for prototyping, for, like, just testing, like, what different interactions would feel like. you can also add audio to it. So it’s a video audio generation model. So you get you essentially can describe both what the visual outcome should be of your click and also what the if there is a sound effect that comes out of it. So we believe that’s gonna be a much more flexible way of building software. Just render. It just, in why generate the code that generates the pixels? Just generate the pixels directly.
Anastasis [00:45:18]: It’s the end-to-end philosophy applying applied to front ends.
Anastasis [00:45:25]: So we think there is a few interesting use case. So you can build creative tools on top of it.
Anastasis [00:45:32]: We think that, for any use case that involves a lot of exploration or, like, educational use case where you wanna learn about a new concept and you want some visualization and like, and open-ended exploration, we think those this is a very powerful, approach. you can imagine new forms of, design, industrial design software that could emerge as a result of those models. And this is all, generated in real-time as well. So, you can build a lot of interesting camera transitions and forms of interaction that are very difficult to build otherwise. And one way in which we evaluate this is what if you try to generate the same interface with Claude by just, prompting Claude, “Here’s an image reference of my interface that I made in Figma or that I created somewhere else. create this particular interaction,” which in this case it’s, drag that object, upwards. and beyond it being slower, it’s also very difficult to capture some interactions by just fully, with just LLMs. So we think that this is likely to be the way that a lot of the future, like, software in the future will be created. and one of the additional benefits is personalization might be a lot easier done with those models. Like, you can essentially try out different prompts based on who is visiting the interface. You can, more easily, prompt engineer the interface to have larger size, text for more accessibility reasons, or you can make this or, like, if you have a particular aesthetic preferences. So we’re very excited about this approach. It’s early days, and I think we’ll need to, make it more cost-effective as well to serve those models ‘cause, running a real-time video model versus just purely rendering HTML, there’s -- the computational needs are much higher. but we do see a lot of potential in this approach to building front-end interfaces.
Swyx [00:47:47]: So we covered this similar thing with Flipbook before with our, Ethan Hara episode with Groq, video. And yeah, I think it’s very engaging visually. I think it’s maybe very good for education, but it’s it does sound expensive. I think there’s an upper bound to how expensive it will be, though, right? Like, the inference cost will go down over time. You’ll figure out ways to optimize it. Effectively, when it pauses, you don’t you’re not receiving human input. You don’t have to generate anything, right? So.
Anastasis [00:48:14]: Yeah, you could also. Like, in this case, you have ambient motion, so there is parts of the screen that might. if you’re let’s say you wanna, visit Paris and then you get this interface that allows you to explore.
Swyx [00:48:29]: People walking. Yeah.
Anastasis [00:48:29]: You have people walking or, like, things happening. But, it’s, it’s a no Yeah, it makes it more expensive because you need to run the model all the time. Maybe you have some looping mechanism so you don’t need to do that. But all those things, I think, is stuff we’ll need to figure out.
Toward a Fully Neural Operating System
Swyx [00:48:44]: Yeah.
Anastasis [00:48:44]: I think our first consideration is let’s make this clearly find some use cases where it’s clearly a much more compelling interaction compared to traditional interfaces. And then it’s a matter of time before it becomes more cost-effective to serve.
Swyx [00:48:58]: Yeah. When it comes to the people walking, I think the approach that makes the most sense to me is Nick.
Anastasis [00:49:04]: Nick.
Swyx [00:49:04]: Oh, God. I keep messing up their name. With Chris Manning and Fanny Yan. I don’t know if you’ve come across them, where they. Mapped to some game engine. I think it’s Unity or something, or Godot. And they you can script some NPC behavior behind that and train on that. Whereas here, you can really imagine whatever you want. Like, that is a UI, right? Like, and it feels, like, more tractable, I guess, to, create a world model of software that is interactable because we have many of examples of that, and you can, do your fancy RL environment stuff on that than it is scaling up to embodied and real-world physical use cases. But this is a nice first step.
Vibhu [00:49:43]: Or, there’s the opposite of you have, like, one B models, three 50 million parameter language models. It just gets so small that they’re just predicting, like, fishes moving.
Swyx [00:49:53]: Small models are now 120 B, so.
Vibhu [00:49:57]: Ultra mini on device.
Vibhu [00:49:58]: But, no, I think it, like, it puts it into perspective, at least the car one for me, like, the applications, right? The amount of work to do that, sure, you only make one model year car per year, but applying this, it’s also a cost-saving to have to manually make all this, right? So it opens up a lot of possibilities, too. I’m curious if you extend this out two, three years, so where do you see things going even further?
Anastasis [00:50:25]: Effectively, the end game of something like interface world models is you have, a fully neural operating system. So I think, Andrej Karpathy has written about that quite a while back. But it’s, You, I think to me it’s, it’s a bit, it’s a bit odd that, we have, for example, with an interaction with an LLM of today, you have this LLM that can talk to you about anything. It can You can take the conversation in any direction. You can It’s very general, so it can solve all those different tasks, but you interact with it through a very rigid interface. And so to me, it’s just a matter of time before the interface itself becomes learnable and becomes, part of the whole loop of, like, you’re not just delivering. You’re delivering an application end-to-end, and that means you’re delivering the language model, but you’re also delivering the render and the pixels and that’s also a learnable component. And the concept of applications might not necessarily. I think we’ll need to figure out new abstractions for software. the concept of application comes from this idea that you need, separate code bases to describe, to, for, to power each individual, tool and each individual application. But you might think of something a lot more unified if you’re. if you have, a video model that’s generating the interface as you go. so it can take context from an LLM and allow you to combine different functionalities that traditionally would live in different applications. So it’s a, it’s a way to solve, software end-to-end, effectively. We also see this as a powerful way to train computer use agents as well. so this is, one way to see this as. And in general, with world models, there is those two directions. One is world models for humans and world models for
Swyx [00:52:24]: Agents
Anastasis [00:52:24]: To train agents.
Swyx [00:52:25]: Yeah.
Anastasis [00:52:25]: And so for every new work of, world models that we do, we have this both uses become possible. So this is a powerful synthetic data generator for training computer use models. It could become, a live, RL environment that you could use to do online RL with a computer use agent, and you can get wide diversity of different interactions, kinds of interfaces, just generated on the fly that, to improve the how robust the, your agent, becomes. So that’s the same also with the world models that we’re working on for a robotics use case as well.
Long Context, Error Accumulation, and Autoregressive Video
Swyx [00:53:02]: Is there a research breakthrough that you’re Waiting for that would unlock the next set of use cases that you really wanna pursue?
Anastasis [00:53:10]: Long context is a very important one, so being able to maintain consistency for long periods of time, and that depends on the use case. So for our characters model, for example, or for the interface world model, it’s easier to maintain long sessions of interaction. If you go into more open-ended worlds that you navigate and you take arbitrary actions in, we, like, there is more the context at which you can and duration which you can generate becomes limited much more quickly.
Swyx [00:53:40]: Yeah.
Anastasis [00:53:40]: So we see more degradation and error accumulation happening. so the biggest challenge with autoregressive models is error accumulation, is you’re feeding generative frames back into the model to generate the next The next frames. And if there is any small errors, they accumulate over time. That’s not a new problem. It’s a problem that LLMs also have, and we’ve seen the ability to generate now really long outputs. So it’s a solved problem, but it’s definitely still a challenge.
Swyx [00:54:08]: Yeah. And what is the state of the art? so for Grok, it would be like 10 to 20 seconds of context going in there for video.
Anastasis [00:54:16]: With our characters models, we’re able to generate up to 30 minutes of video autoregressively.
Swyx [00:54:21]: Yeah. But that’s just for the avatars.
Anastasis [00:54:24]: Yeah. So if we look at, GWM Worlds, which is more our open-ended world exploration model, it’s, it’s on the order of a few minutes, which is Yeah, so
Swyx [00:54:35]: Probably enough for people because you have to cut to the next scene anyway, right?
Anastasis [00:54:40]: Yeah, it’s not, it’s not the ideal game experience if you have to restart every few minutes. So I think. But, I think it’s. Yeah, for certain kinds of game experiences, you can work around it. ideally, you are able to just generate forever, and it doesn’t, it doesn’t degrade. And I think that’s a matter of time before we get there.
Swyx [00:54:59]: Yeah. Genie has, like, one, max one minute?
Anastasis [00:55:01]: Right. Yeah.
Vibhu [00:55:02]: This was your. You did a study on robotics. I think I also have just your Runway Robotics page, though. Is this better?
GWM Robotics and Sim-to-Real Evaluation
Anastasis [00:55:11]: So last year we released Gen-4.5, so that was our latest base model. We’ve been As I mentioned, we’ve been doing all this work in world models, and which essentially a lot of our approach to world models is how do you take a bidirectional diffusion model and make it autoregressive and make it accept actions? So instead of being a video you watch, it becomes a simulation that you step in, and you can, control it every step of the way. You can explore counterfactuals, like what happens if I take this action versus if I take this action. And GWM-1 was the it’s the world model that we built on top of Gen-4.5. So we did all this autoregressive and like, distillation, auto-regressive and then step distillation on top of Gen-4.5. And one of the biggest use case that we saw for GWM-1 was in robotics. One thing we like to say is we as we scaled video models, we accidentally, created one a state-of-the-art model for robotics, by just scaling video models. So we realized at some point, mid last year that robotics labs that are coming up to us and asking to use video models for synthetic data, asking us to post-train our video models to work really well for robotics, so that they can use that to generate variations. That was the first use case that we saw. And then increasingly became clear that the models will be useful beyond just creating synthetic data to train robotic policies. They would also be very useful as simulators. So that means that you can use, a video model online to test how your robotic action model performs. So you can take an action role and then get the outcome of the action inside the world model and then continue that loop like this closed loop simulation. And you can use that to evaluate how well your robotics model works. and the biggest thing that I think you need to solve if you want to build a simulator is establishing real-world correlation that if you take an action inside the world model, if you take the same action in the real-world, you get a similar outcome. So that was the goal of some work that we did earlier this year. So if you go to the first link. So that was, essentially wanted to establish that, real to sim correlation for our world model, so that if you do a series of actions inside the world model and if you do the same actions in the real-world, you get similar outcomes. And we took our GWM-1 model and we used some benchmark data that there is this Roborina, benchmark that’s very commonly used to evaluate how well do different action models perform. And we use the same scenarios and settings and embodiments inside our world model, and we measure the correlation of how well did the action model perform inside the world model versus in the real-world. And we saw that we could get very good correlation between our world model and reality. And that means that if you want to evaluate how well your robotic policies perform, you can scale that much faster inside simulation instead of having to do that with actual physical hardware. And so that was a first indication that our models could be, quite useful in robotics. And we saw as we were working with robotics labs that became like the first use case where they could use video models in a way that feed into their training pipeline.
Vibhu [00:58:40]: Can I ask what
Anastasis [00:58:41]: Yeah
Vibhu [00:58:41]: The difference was from four point five to solving that? So the sim to real gap has always been the issue, right? You train a robotics model on video data, it doesn’t generalize to real-world, and the simulation had an issue. So seems like you solved it, but how?
Anastasis [00:58:56]: Yeah. So a big problem with simulators is, if you’re trying to simulate rigid objects, like it works quite well if you can describe the physics of objects very accurately, then you’re able to use, Isaac Sim or MuJoCo or one of the traditional simulators. But for more complex interactions with cloth, for example, or, like slippery surfaces, with the all the complexity that you want to be able to solve with the manipulation, with an action model that solves manipulation tasks, it’s very difficult and so time-consuming to build, for each of those environments and each of those tasks, build the simulated version of that, the digital twin of that environment. Whereas with a world model, you just need to provide the first frame and then you just can roll out the policy inside the first frame. So whereas, we compare it to methods that required like 3D scanning an environment and then 3D scanning each individual object before you can now, you can bring that to simulation. whereas with a world model, you just take a picture of the environment and then you’re able to test how your policy performs. Our general thesis on robotics is, there is companies that are leveraging a lot of teleoperation data to train robotics action models. There is now companies that are using, humie data, which is, essentially human, egocentric video where humans use robotic creepers to perform different manipulation tasks. And then there is companies that are focusing on egocentric data, which is, you strap a GoPro on someone’s head and then you capture them performing a task. We think that, and all those are great source of data for training robotics models, but the most plentiful source of video data is third-person video data. It’s And if How do we as humans learn how to perform different tasks? A lot of it is by observing others perform those tasks. We don’t learn from first person. We do some trial and error and like, to learn different things, but. Ultimately, a lot of what we learn how to do in the world, we learn by watching other people do it. And that’s how when you’re pre-training a video model, you’re essentially doing that. It’s a lot of third-person video footage of people performing different tasks in the world, people doing sports, people doing household tasks. And our main thesis is that video pre-training, once you do that, you can then adapt a model to be useful in robotics use cases with way fewer hours of actual robotic data. So you require way less teleoperation data, which is very difficult to scale. and even if you look at egocentric data, which is a bit more easy to scale compared to teleoperation data, which requires actual hardware,
Why Third-Person Video Is a Powerful Robotics Pretraining Source
Anastasis [01:01:55]: It’s still three hours of magnitude less of that exists in the world compared to third-person video data out there. And so our thesis is and generally, like the most plentiful source of data will ultimately wins. Third-person video data pre-training is the right starting point for models that, you want them to generalize and be able to deal with new environments, new tasks, things that you haven’t seen during training. That’s the motivation for why we think our models are especially useful in robotics, settings, and we’ve seen that to be the case, as well.
Swyx [01:02:32]: You said pre-training. So maybe it’s like third-person pre-training, first-person SFT? Is there like a curriculum that you can introduce?
Anastasis [01:02:41]: Exactly. So if we look at GWM Worlds, so GW so GWM Robotics. So digitally in robotics, it starts from Gen-4.5.
Vibhu [01:02:49]: It’s the same video diffusion backbone, right?
Post-Training World Models for Robotics Embodiments
Anastasis [01:02:53]: Exactly, yeah. So you start from the base video model, the one you’re using to generate, cats and dogs and other interesting stuff, and then you, fine-tune on a very small number of hours of robotic data. So it’s something on the order of hundreds of hours compared to if you were to pre-train a robotics model. The current pre-trainings go up to, a hundred thousand or like millions of hours of data. And you’re able to get quite good performance, quickly, because the model leverages all the things that it has learned about the world, physics and human dynamics and the tasks that people care about from pre-training. And ultimately, you want those models to generalize. You don’t want to just be able to perform the tasks that it has been doing training. And the diversity of actions and environments that you have with a pre-training video dataset is much larger than, what you can realistically capture manually.
Vibhu [01:03:54]: How is the scale looking like for the post-training? Like, do you still wanna do, is it like roughly ninety percent of the compute in regular video diffusion model and then scale up a lot, or do it like we want different robotic models for different tasks, or just the one base really good world model can also apply to robotics?
Anastasis [01:04:14]: So currently, we are post-training our models for specific, embodiments that we for particular partners. So if they have a particular single-arm robot or a bimanual robot or a humanoid robot, we would post-train our GWM robotics model on their particular dataset. Over time, we see the different variants of GWM unifying. Like, I would expect, if a year from now or two years from now, you have a single world model that can simulate manipulation tasks, it can simulate navigation, which is a lot of the gaming world models are navigational world models. You’re moving around the space, and it will also simulate human behavior. So that’s the character models. So instead of having three different models, you have a single model that’s able to. ideally, you’re able to simulate what it’s like to be in the world. You’re moving around an environment. You’re maybe performing different tasks. you’re talking to other people. And that happens with, the same, a single real-time video model that’s generating that.
Vibhu [01:05:17]: Do you think you can solve self-driving? So if you are learning to drive a car in a simulator, you have a world model. Your robot is car can manipulate so many axes. How far off are you from something like that?
World Action Models, Self-Driving, and Learned Policies
Anastasis [01:05:31]: So world models
Vibhu [01:05:32]: Or a really good ADAS system?
Anastasis [01:05:33]: World models are definitely being applied to, self-driving, research right now, mainly for evaluation use cases, but our focus has been more on robotic manipulation. We’ve done some work on AV, world models as well. but yeah, we do think that world models are and video models are the best starting point for both simulators and also policy and the action models. So that’s, that’s the other side to this, is that once you have a great world model, then you can just add an action head, and it can predict actions as well. One way to think about it is if you take the starting frame of a scene with a robotic arm and you ask, you prompt the model, generate the arm picking up an object, it would And if it generates an accurate enough video, then it should also be able to generate the exact poses, in 3D that the arm should take to perform the same action. So this is the direction that’s now the popular term for it is world action models, which is you’re starting from a video model, and then you’re adding an action head to predict the actions, and it becomes a policy, essentially.
Swyx [01:06:43]: One thing I’m also impressed by is how much data you need to train these kinds of models. You probably can’t say exactly how much, but like, the original, diffusion models, and from what I know, even of the open source Chinese models, it’s not that much data. Isn’t it surprising?
Anastasis [01:07:02]: What do you define as much data?
Swyx [01:07:05]: Yeah, and it just comes, goes in. Is the token count still relevant?
Anastasis [01:07:09]: So it’s a bit more complicated and,
Swyx [01:07:11]: What is just gigabytes, right?
Anastasis [01:07:13]: Yeah, hours of video, right?
Swyx [01:07:15]: Yeah. Yeah. I feel like something that’s interesting is it seems like the, let’s call it tokens to param counts in language models has really, maybe they’re three years ahead or whatever, seems to be a lot higher than, video models still, even though technically video has more information, per bit. I don’t know if it seems intuitive or maybe there’s just a lot of, like the variability between a pixel to the next pixel is not that high. So, like, maybe there’s just a lot of information that is repeated.
Scaling Video Data and the Lucid Dream Test
Anastasis [01:07:47]: My answer would be it’s still very early. Like, the training video models will scale way further than it
Swyx [01:07:55]: Yeah
Anastasis [01:07:55]: Currently is, and you’ll have capabilities that go much further than the current models can do. So one thought experiment that, I like to use, it’s, it’s almost like the Turing test of video models or like the Turing test of world models, go, I call it the lucid dream test. It’s you have a
Swyx [01:08:14]: You mean the actual person lucid dream?
Anastasis [01:08:17]: It comes from this idea
Swyx [01:08:18]: Lucid rains, right?
Vibhu [01:08:19]: Lucid dreams is telling you’re dreaming while you’re
Swyx [01:08:22]: Yeah.
Anastasis [01:08:23]: Yeah, exactly. So lucid dreaming is when you realize you’re
Swyx [01:08:25]: In a dream
Anastasis [01:08:26]: Inside a dream, and then you
Vibhu [01:08:28]: Play around
Anastasis [01:08:28]: Be able to control what happens in
Swyx [01:08:30]: No, there’s also an inference guy called Lucid Rains. Yeah. Or quantization
Anastasis [01:08:33]: Very prolific, person. Yeah. So let’s say you have a VR headset and you’re in a room with and you’re wearing a VR headset, and that VR headset, most of today’s VR headsets have a pass-through mode, so you can see directly what’s in front of you in the world, or you can render something inside the VR headset. And there’s gonna be a point where those interactive real-time video models become good enough where you wear the headset and you’re in the same room and you’re walking around and you’re kinda and you’re interacting with objects. You’re able to move freely in that room and do, and interact with any object. And at the end, someone asks you, “Did you were you using pass-through mode, or were you -- or was this, rendered or generated, footage?” And if you cannot tell for sure if that was what you were seeing as you were interacting with and moving around the world was generated or it was, pass-through mode and was just what was happening in front of you, that’s an indication that the models have become good enough. And we’re not, we’re not close to that yet. And a lot of it is just this idea of really simulating dynamics and counterfactuals well. Like, if you ask a video model to generate a person scoring a goal versus a person failing to score a goal, it would do a better job at scoring the goal because there is a bias from the training distribution. There is a lot more videos of the person succeeding at scoring the goal. But if you have an interactive model, you want it to be able to generate counterfactuals. Like, if I take this action versus this action, you want it to generate equally realistic outcomes. so that’s, I think, the big gap between video models and world models is that idea of the counterfactual generation. And if you want a great model for robotics, you wanna simulate failure very well, because whether you’re using it for evaluation or you’re using it as a in an online RL loop in the future, you wanna be able to have the model try and fail to do things and improve. and so in order to do that, you need to be able to simulate things failing.
Swyx [01:10:42]: This is the only domain where you have too many successful examples and not enough bad examples. Should be easy to generate failure.
Vibhu [01:10:51]: Oddly enough, I think, like, early image video models weren’t good at being human realistic, right? Like, you see aa lot of the high-res 4K, like, professional photography, but not just everyday life, like normal picture, right? Everything looks like it’s professionally generated, like professional pictures, but not just like normal, like, messy cables on a desk.
Swyx [01:11:13]: Okay, so there’s, there’s this stuff. one thing we also covered that you guys have, video agents that you launched. I guess, how does the traditional, let’s call it frontier, like, autoregressive LLMs, like, feed in, to all this? They’re driving ro your robotics models, or are they driving others, your video agents, production, anything where you see the overlap of autoregressive and diffusion, let’s call it?
Counterfactuals, Failure Data, and World Model Evaluation
Anastasis [01:11:41]: Yeah, so harnesses are really important across all those different use cases. So we have this video agent, which is essentially an LLM that is very effective at tool use of different, image models, video models, and helps you through creating a project end-to-end. So, very often in, like, a traditional advertising flow, you have a brief, you start from it, and then you generate some a storyboard, and then you generate the video. A video agent and, or runway agent helps you through that whole process, and it helps you also analyze performance data. For example, how well did this ad perform versus this ad, and then generate me more of the based on those learnings, figure out what to generate. We think that the harness is a very important piece of the pipeline. as I mentioned, all the video production, all the production video models use some prompt completion that happens, and we expect, that to become more and more complex and more, you generate longer and more detailed descriptions before you use the diffusion transformer. I do think eventually, there’s increasingly this unification into omni models where you have the you’re training the models end-to-end to both do autoregressive text prediction and also, diffusion as well. So you’re predicting the next token, of like you’re, you’re maybe using some reasoning and planning of the scene, and then you’re passing it into the diffusion head that’s generating the pixels.
Video Agents, Harnesses, and Omni Models
Swyx [01:13:10]: Yeah. I think currently maybe only Gemini and Qwen do it. I-I’m not sure which of the Chinese models are omni, but yeah, it’s, it’s not, it’s not a very well, popularized modality, I guess.
Vibhu [01:13:25]: It’s an interesting use case when you think about it, right? Because not only do you have to end at like language model reason, diffusion had generate, you don’t have to output there. You can go back in to feed that output to the same model, reason again on improvements, and it can do a lot of loops just in its own. I guess the question is like, do we need that or can we just do agent scaffold, like do it outside the model? Is there a big benefit to doing it in?
Anastasis [01:13:54]: I think there’s generally the trend of something is first done by a harness and then it becomes part of the model, right? So you had the chain of thought prompting where you had to do this super detailed system prompts to
Swyx [01:14:07]: Yeah, step by step
Anastasis [01:14:08]: Get the output. And now the model generates the reasoning trace by itself before it gives you an answer. And in the, in video models similarly, a lot of the video models of the early days were single-shot video models, and you had to use some orchestrator to turn, generate multiple shots in parallel, and then turn it into an actual video.
Swyx [01:14:28]: Or in ComfyUI, just all over the, all these nodes.
Anastasis [01:14:31]: Yeah, like a spaghetti workflow. and now you have multi-shot video generation where you have the you directly generate multiple shots. And there is a benefit to that because then the video model learns some. to generate a single shot well, you need to figure out a lot of stuff about the world. to generate multi-shot video well, you also need to get some, like, video editing instincts. Like, you need to figure out what is the right pacing of shots. And also, LLMs are not that good at it. Like, they’re not that great video editors. If you ask a LLM to take some videos and then auto-create a edited video out of that, it would feel uncanny. So I don’t think LLMs are that good yet at being video editors. And I think there’s benefit to learning that end-to-end. so I would expect, the training generally is the things that, you need the harness for eventually get injected into the model itself, and you learn that end-to-end.
From Harnesses to End-to-End Learned Video Editing
Swyx [01:15:33]: Do you find that you need to hire engineers who can. or researchers who are also artists to infuse that taste, or do you have artists in residence to distill them?
Anastasis [01:15:44]: We have a large creative team that’s very actively involved in the, in training those models, like on the, in every part of the way. And like, how do you caption video as well so that you capture the stuff that you need for, like, the cinematography, the aesthetics, the camera direction in as detailed ways as possible so that you’re able at inference time to elicit that through the model? we have our creative team also does a lot of evaluation of like, what constitutes a usable video out of those models. And so they’re very involved through every part of the process. And I think that’s one of the special things of Runway is just that mix between like creatives and researchers sitting by, side by side and working together to build the next generation of our models. I think that’s been a really important piece to, how we’ve operated as a company.
Swyx [01:16:37]: Yeah. In some senses, though, you can only do this in New York.
Vibhu [01:16:40]: It’s
Swyx [01:16:40]: Maybe, you have other offices, but like, I try to find some poetic, significance in the fact that you are a big New York company.
Anastasis [01:16:49]: As there’s a few parts to being New York. there is that intersection of all those different industries and, like, media, advertising, like
Swyx [01:16:57]: Yeah, this is very advertising.
Anastasis [01:16:59]: The, like the art scene is New York. Not to say anything bad about San Francisco, but, it’s. There is more going on. There is that component, and there’s also, I think we benefit from being outsiders and thinking of things a bit differently, like not being in the same, like, hive mind of,
Swyx [01:17:19]: BВС
Anastasis [01:17:19]: ASI, of Bay Area and, like, taking. and also taking our time to get where we are today. Like, building the, growing the team intentionally and bringing people who are, yeah, both on the creative side and also on the engineering research side. There’s huge talent pool of amazing people in New York, so that hasn’t really been a problem.
Creative Taste, Artist Feedback, and Runway’s New York Advantage
Swyx [01:17:41]: Congrats on everything. what are you hiring for? what should people look forward to, for the future of Runway?
Anastasis [01:17:49]: We’re hiring across the board. I think this is probably the most open roles we’ve ever had in the history of Runway. we’re growing our research team quite significantly. So if you’re, if you’re excited about video models, if you’re excited about world models, if you’re excited especially about robotics, the robotics team, we’re hiring roles in the robotics across, software, hardware, and research. so definitely reach out.
Swyx [01:18:13]: And, a lot of people don’t have direct robotics background, but what should they have, if they want to be useful in robotics?
Anastasis [01:18:21]: So ideally, some experience with learned policies, would be
Swyx [01:18:26]: Just RLs
Anastasis [01:18:27]: Good for robotics. but we tend to hire generalists as a philosophy and, like, people who learn really quickly. but some experience in the, in domain expertise in robotics is something that we’re, we’re definitely looking for the next months. and then we’re scaling the go-to-market team significantly. There is, a wide, like, very active enterprise adoption happening around video models at the moment, and, we’re really trying to, respond to all the demand.
Swyx [01:19:00]: Yeah. Great. You wanna talk about the, open source robotics stuff?
Vibhu [01:19:04]: Sure. It was just random notes we had.
Vibhu [01:19:07]: NVIDIA launched Cosmo. I guess it’s interesting. So, you’re a founding member AI labs to build open source world models in physical AI. - Anything else to talk on here is open research?
Anastasis [01:19:20]: The biggest thing is that, as I mentioned, while models are still, early, like there is still so much that we you can scale and those models further, so much more advancements and things that we can figure out and how to improve those models further. And I think this is, it’s important that some of this research happens in the open and figuring out what is some incentives for different companies to come together to bring some of that research into the open and open source. And so Cosmos Coalition was a initiative that we co-founded with NVIDIA to bring some of that research as open source. And that could mean open weight model releases. It could mean benchmarks that measure physics and things that people care about when building world models. It could mean infrastructure. So really, how do we grow the ecosystem of world models and make that something that also it’s easier for a developer, a researcher that’s just starting out that is excited about world models to contribute to the field.
Hiring, Robotics, and Enterprise Adoption
Swyx [01:20:19]: I think it’s a there’s some amount of like, is this also our response against the Chinese world models that are being released, or is there not part of the consideration?
Anastasis [01:20:29]: I do think it’s, it’s important for NVIDIA models, if you look at the leaderboards of video models, I would say right now the majority of models at the top ten, top twenty are Chinese models. There is, only a handful of companies that are made it to the leaderboard from like the US or the West.
Swyx [01:20:50]: Yeah. We’re doing better with images, but with video we’re very behind, right?
Anastasis [01:20:53]: And so I think it’s definitely important that we invest more broadly as a community to make sure that we can those models can we have competitive models
Swyx [01:21:02]: Yeah
Anastasis [01:21:02]: Out there.
Swyx [01:21:03]: But like what’s to stop us from just distilling from them?
Anastasis [01:21:06]: I don’t know if that’s the best long-term
Swyx [01:21:08]: Not gonna mention that they won’t
Anastasis [01:21:09]: That you’re bounded by the performance that you can. It’s, it’s almost a bit of a pessimistic
Cosmos Coalition and Open World Model Research
Swyx [01:21:14]: Like
Anastasis [01:21:14]: View that you can get better. you can
Swyx [01:21:17]: It’s free data. it’s, you might as well. Like if they’re, they’re doing it for like, for the text language side, they might as well do it for the video side the other way.
Anastasis [01:21:25]: Yeah, I do think we’re, we’re quite capable of training great models
Swyx [01:21:29]: Okay
Anastasis [01:21:30]: Without distillation at the moment. Yeah.
Swyx [01:21:32]: Yeah.
Vibhu [01:21:32]: So anything you have to say on benchmarks and evals? Like, I feel like what I’m hearing is a lot of people really like arenas for video and image models, customers and whatnot as well. They only want the best on the leaderboard, and they refer to arenas a lot more than language models seem to do. But any notes on benchmarks, what’s lacking? How does the average person compare while these both look really hyper-realistic? More than that, outside of we did talk about like robotic simulation, the physics and all that, but anything to say?
Anastasis [01:22:05]: I think it’s the opposite in some ways. I think people, generally creatives and artists and marketers, other like people that are using our platforms, I think rely less on, arena scores. And it’s, it’s just so easy to, generate with a bunch of different models and then compare the results visually. Like one nice thing about image and video models is you can immediately tell with your eyes like what feels good from an aesthetic standpoint. Like any artifacts, any issues with the physics of those models, you can immediately tell. and so that’s it’s easier, I would say, to evaluate, as a human. there is also those models than it is in language models where you have those very complex math and coding and, tests where it becomes a lot more harder, I think, for humans to evaluate and can discriminate between the performance of models at a time. So I think in practice, people just test out the same prompt with a bunch of different models and see what the results look like. And right now in Runway, you can use our models and you can use third-party models as well. So it’s, it’s very easy to do that.
Benchmarks, Arenas, and How Creatives Evaluate Models
Swyx [01:23:12]: Amazing. We’re gonna end with the AI Runway AI Summit. The last societal issue, I guess, I don’t know if this is a thing, is the, you are at the tension between artists and creatives and AI. A lot of people in that community hate AI. the people that are in the Runway community don’t mind using tools. it’s just another brush. But, how have you seen the sentiment change?
Anastasis [01:23:37]: Our perspective, yes, it’s just another branch, brush. It’s just another camera. It’s, it’s the latest of a long generation of tools.
Swyx [01:23:46]: Technology in art.
Anastasis [01:23:47]: Technology.
Swyx [01:23:47]: Yeah.
Anastasis [01:23:47]: And art and technology have evolved together. I think there’s been a pretty significant shift over the past few months, and it came. some of it you can see with a lot of public figures speaking out in favor of AI and being, like in Cannes, you saw a few directors speaking in favor of AI. We had Ron Howard in our film festival. There is, Mark Scorsese also adopting AI models. So you have more of those stories coming out every day of like a well-known figure, speaking in favor of AI. And it’s just a matter of, in my mind, it’s those models are becoming more and more demystified. I would say I have also a bit of a hot take that one of the things that made the initial response to those models maybe a bit more heated than it needed to be was this idea of text to video of, you have a single text description and you get back a -full video.
Artists, AI, and the Evolution of Creative Workflows
Anastasis [01:24:48]: Yeah, there was a misconception. you can generate it to our feature-length film, but the models of today now take a lot of references. They take they are very controllable. And I think when people see a tool that allows, affords many degrees of freedom and control, they respond to it differently. And it matters less that it’s a generative model than the fact that you can steer it to the direction that you want. and so. I think when people look at, complex workflows on top of those models, when they look at, all the ways in which you can steer them and you can provide now with some of the latest models up to fifty references, like the conversation becomes a bit different because it feels much more like a
Swyx [01:25:34]: Storyboard
Anastasis [01:25:35]: A tool
Swyx [01:25:35]: Yeah
Anastasis [01:25:35]: Versus, like, something that a magical entity that figures out, like, the, your entire film for you.
Vibhu [01:25:43]: Any notes on, like, workflows changing for people in the field? Like, I think engineering at least has had a lot of people where they’re like expectations have changed. I’m, ten X, a hundred X more productive, and you can get a lot more done. same thing as, you’re making dev tools for creatives. any notes there? Like, there’s some people that don’t wanna adopt, some that do. Like, anything?
Anastasis [01:26:08]: Yeah. So I think, in terms of, like, what people care about, I see that we have gone through a few stages. So we started from a stage where the main thing that people were looking for was quality. Like, as, we scale those models, the quality improved dramatically. That’s something that people still care about, but it’s, it’s now in addition to controllability, like being able to steer those models with references, with, different kinds of inputs, with storyboards. And now my sense is increasingly people are gonna care about latency more and more. As those models become better, the ability to iterate very quickly becomes more important. And, like, if you can, with a single prompt generate ten different, outputs, like, almost instantly, you can explore way faster than before. And you get some of the magic that characterized the creative tools of the past, like Photoshop was instant. and we lost some of that with generative models. You’re waiting for two minutes to get back a video, and I think we’re gonna bring, some of that back now with the
Vibhu [01:27:08]: Real-time
Anastasis [01:27:08]: Real-time models.
Vibhu [01:27:09]: Yeah. Exciting. And
Latency, Real-Time Generation, and the Future of Creative Tools
Swyx [01:27:11]: Exciting. the last thing we’ll plug is this one, Runway
Vibhu [01:27:14]: Summit
Swyx [01:27:14]: Summit. You’re finally doing this in SF?
Anastasis [01:27:18]: Yeah. So, we’re very excited about this. So this is, in late September thirtieth, we’re doing a summit on, primarily focused on physically high and real-time video generation. We have panelists from NVIDIA, Physical Intelligence, Botco, DeepMind. Yeah, it’s gonna be, I think, a very interesting series of conversations. We try to make the panels really technical and, elicit actual substantive discussion and hopefully some interesting disagreements and interesting debates on things. And, yeah, the there’s tickets available. Hope people can join.
Swyx [01:27:56]: Since you mentioned it, what disagreements and debates should people think about, or do you expect?
Anastasis [01:28:04]: So it’s things like, there is, one debate right now in the robotics world is, VLA’s versus world action models.
Runway AI Summit and the Big World Model Debates
Swyx [01:28:11]: Okay.
Anastasis [01:28:11]: So there is labs that are really betting on one of those two directions. there is like what is the best source of data to train robotics models?
Swyx [01:28:21]: There’s just the third-party, first-party that we talked about.
Anastasis [01:28:24]: Yeah. There is, the people who really believe in further scaling teleop data versus leveraging more large-scale video data. So that, those are some of the. And then there is, the world models debates of predict pixels directly versus something like JEPA versus a more 3D-based, 3D-based approach. so I think we’re at a nice time in world models because there is still that active debate happening on, like, what is the best long-term direction. I feel very strongly that it’s video predict pixels directly and scaling video generation models is the right approach. But it’s, I think there is a lot of interesting, debate happening, by researchers on, like, what is the best path to take.
Swyx [01:29:09]: It’s interesting that it’s all on, like, let’s call it the policy layer and the data model layer. Is the physical side is completely solved? Like, all the sensors, all the actuators, all these things are. We have everything that we need?
Anastasis [01:29:23]: I don’t think that’s, solved either.
Anastasis [01:29:25]: It’s definitely,
Vibhu [01:29:27]: Different problems.
Swyx [01:29:28]: It’s, it’s like
Anastasis [01:29:29]: Yeah
Swyx [01:29:29]: I wanna dream about all these things, and then I get, I buy a robot or I buy, I try to assemble my own, and I can’t even get the motors to, like, work right. Right? Like, and it’s you’re dealing with very sensitive, equipment that has, voltage and power and, like, heat and all these things which, you, abstracted away. We’re sitting here, we’re talking about software and talking about models, but, like, really you have to deal with those kinds of things too.
Anastasis [01:29:56]: Yeah. And, I think I’m, I’m, I’m generally also not opposed to incorporating other modalities into our models like we’ve seen.
Multimodality, ImageBind, and the Maximalist World Model
Swyx [01:30:04]: Yes.
Anastasis [01:30:05]: The simplest case is they can generate video and audio at the same time. So they can generate RGB, and they can also generate, they can generate sound and audio. But my. I’ve written about this as like what does the maximalist version of a world model look like is you’re incorporating more and more modalities from the universe And you’re training a model on different scales of observations as well.
Swyx [01:30:28]: X-rays.
Anastasis [01:30:29]: And so, yeah,
Vibhu [01:30:30]: You got a good essay that people should read on
Anastasis [01:30:33]: Yeah. Yeah
Vibhu [01:30:33]: Real-world.
Swyx [01:30:33]: No, Meta released a model that was, like, six modalities in one, right?
Vibhu [01:30:37]: Yeah.
Swyx [01:30:37]: I forget what the name of the thing was, but it was like, yeah, okay, depth is one of them, but depth is like a transformation of RGB in some sense.
Vibhu [01:30:45]: ImageBind.
Swyx [01:30:45]: ImageBind, yeah.
Vibhu [01:30:45]: Yeah.
Swyx [01:30:46]: What other modalities? They had heat?
Vibhu [01:30:47]: Audio, depth, heat, text,
Swyx [01:30:51]: Whatever IMU is.
Swyx [01:30:52]: I do think, like, you might as well do ultraviolet. You might as well do, like, just whatever other modality you feel like, ‘cause it’s all data to the model.
Anastasis [01:31:01]: Yeah. And, a big bet is also that there is transfer between all those modalities.
Swyx [01:31:05]: Yeah. Yeah.
Anastasis [01:31:05]: So one of my favorite, examples, which is quite old at this point, is there was this fine-tune of, Stable Diffusion that was called Riffusion Which was
Swyx [01:31:14]: The music one. Yeah.
Anastasis [01:31:15]: Yeah, just fine-tuning, Stable Diffusion on spectrograms.
Swyx [01:31:18]: Spectrograms.
Anastasis [01:31:18]: And it became a quite capable music generator. Right? So there is probably Spatial patterns, so like spatial-temporal patterns if we’re talking about video that emerge at different scales and different modalities. And so there is some degree of, meta-learning that the model has done that allows it to learn faster if you start from a just a model trained on images and train it to predict audio than if you train from scratch on just audio. and there is some other interesting examples. So there is this project called The Well. It’s, it’s a dataset of physics and numerical simulations in physics and biology and a bunch of other domains. So it’s, so it’s essentially different physical systems across very different scales of space and time, from like astrophysics to low-level like atomistic interactions. And we’ve seen. we’ve done some work on this, and we’ve seen that we can take our video model where, real-world video looks nothing like this, and you can fine-tune it on those numerical simulations and just treat them as RGB frames. And you get reasonable performance much quicker than if you just train from scratch.
Anastasis [01:32:36]: Yeah.
Vibhu [01:32:36]: I think we’ve seen this across languages where
Swyx [01:32:38]: Yeah, DeepSeek-OCR as well.
Vibhu [01:32:40]: Yeah, DeepSeek-OCR.
Swyx [01:32:41]: Like, you don’t have to tokenize text. Like, you can just throw them in as images.
Vibhu [01:32:44]: There’s a lot that happens in that base pre-training. Like, there was an argument a long time ago of people saying, “Oh, humans have so many, sensory representations, right? Smell, touch.” Models have a whole two more modalities that we’ll like, that we don’t even have data for. And it’s like, okay, you take AQI sensor, like you can try this stuff, but there’s so much happening in just the base trainer on that you don’t get as much from these little things.
Scientific Data, Cross-Modal Transfer, and Omni Models
Anastasis [01:33:10]: Yeah, exactly. And I think that’s what it solves is data scarcity.
Vibhu [01:33:13]: Yeah.
Anastasis [01:33:13]: So you don’t have as much. You have so much video data available, but you don’t have, like olfactory data that
Vibhu [01:33:21]: The cool thing is it goes the other way too, right? So if you wanna do physics, like if you wanna measure this or you wanna have a diffusion model do audio, it transfers really well. So like in your case, the little bit of post-training for robotics gets a video model to use its fundamentals in another domain. So we can apply that to other stuff too.
Anastasis [01:33:40]: Yeah. And if we look at, like how do you make those models more useful for in scientific domains, and if you look at AlphaFold, they’ve had all these very. Because of the data, the limited amount of data that it needs to be trained on, it’s it’s very fine-tuned architecture just to solve, protein structure prediction. But if you take all those disparate sources of scientific data and you bring them together under a single model, like I think that’s an approach that can help us solve new kinds of problems across science by leveraging all the learnings from one modality or one set of, data to another. So very early days for that direction, but I do think that’s where ultimately what the end game of simulating the world is. You’re not just using RGB. You’re using RGB as a starting point, but you can incorporate more and more modalities of the universe and leverage the transfer that happens from learning from one to the other.
Vibhu [01:34:42]: I guess the follow-up there is what’s the drawback of omni? Like, why is everything not an omni model? Also, why not now, and why. Would you start from language backbone or image video backbone and then go omni from there? Does it matter?
Anastasis [01:34:57]: Yeah. We need to take it one step. We need to solve robotics first, and then we can go into
Swyx [01:35:02]: Solve everything now.
Anastasis [01:35:05]: Yeah. I do think there is a lot of open-ended research that needs to happen for, those omni models. There is a lot of things that require careful consideration when you’re bringing multiple modalities into a single model to predict. But I think, I expect those to be solvable.
Closing: Film Festivals and the Future of AI Video
Swyx [01:35:23]: Wonderful. you’ve been very generous with your time. Congrats on all your success, and, yeah, I’m excited for the, AI Summit, or physical AI Summit.
Anastasis [01:35:32]: Yeah, thanks for having me.
Swyx [01:35:33]: And yeah, and people should check out the film festival if it’s in town, right?
Anastasis [01:35:37]: Yeah.
Swyx [01:35:37]: Yeah. You’ll be gonna be touring all over the place.
Anastasis [01:35:39]: Yeah. Next year we’re probably gonna do that. So we do film festivals every May or June of
Swyx [01:35:45]: Yeah.
Anastasis [01:35:45]: And we did the last one in New York, LA, Tokyo, and at the AI Engineer,
Swyx [01:35:52]: Yeah
Anastasis [01:35:53]: Fair.
Swyx [01:35:53]: Yeah. Yeah.
Anastasis [01:35:54]: So yeah, hopefully even more places next year.
Swyx [01:35:57]: No, I think like someday, you will be hosting the Oscars of AI video, and, I think people should like take this very seriously as like a potential career they can have.
Anastasis [01:36:07]: The Oscars of AI video will be called the Oscars.
Swyx [01:36:10]: All right. All right. Thank you.
Anastasis [01:36:14]: Thank you.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Plus de podcasts Business
Podcasts tendance de Business
À propos de Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0.
We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al.
Full show notes always on https://latent.space
Sponsorship and business inquiries: business@latent.space www.latent.space
Site web du podcastÉcoutez Latent Space: The AI Engineer Podcast, BFM Bourse ou d'autres podcasts du monde entier - avec l'app de radio.fr

Obtenez l’app radio.fr gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités
Obtenez l’app radio.fr gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités


Latent Space: The AI Engineer Podcast
Scannez le code,
Téléchargez l’app,
Écoutez.
Téléchargez l’app,
Écoutez.



























