70 épisodes
- AI spent the last few years competing on intelligence. Now the next major AI race is about speed.
As AI shifts from simple chatbots to agents that reason, write code, call tools, search the web, and coordinate other agents, latency compounds. A task that requires dozens or hundreds of sequential model calls can quickly turn seconds into minutes — or even hours. That makes inference speed a fundamentally different problem in the agentic AI era.
In this episode of So What About AI Agents?, Philippe Trounev sits down with Vasanth Mohan of SambaNova Systems to unpack what actually makes AI faster — from model architecture and memory bandwidth to batching, specialized accelerators, and next-generation inference hardware.
Vasanth explains the two major stages of inference — prefill and decode — and why they create very different hardware bottlenecks. They discuss operator fusion, parallelization across chips, memory bandwidth, and why reducing latency often comes with a significant cost tradeoff.
They also explore why coding agents are one of the clearest use cases for premium inference, how AI providers may eventually offer commodity, premium, and ultra-fast tiers of tokens, and why businesses willing to pay for speed may gain a meaningful productivity advantage.
Vasanth also shares early performance figures for SambaNova's upcoming SN50 architecture, including benchmark results around 800 tokens per second on a large model, compared with roughly 300–400 tokens per second on GPUs in the cited comparison.
We also get into a slightly crazier question: are AI agents already helping engineers design the next generation of AI hardware? The answer is increasingly yes — although humans are still very much in the loop.
In this episode:
Why AI agents make inference speed dramatically more important
Why sequential agent workflows create a latency bottleneck
Prefill vs. decode explained
GPUs vs. specialized AI accelerators
The relationship between speed, batching, throughput, and cost
Why memory bandwidth matters for large AI models
What SambaNova's SN40 and SN50 architectures are designed to solve
800-token-per-second AI inference
Why coding agents benefit so much from faster models
Commodity vs. premium vs. ultra-fast AI inference
Whether faster AI becomes a competitive advantage
How AI agents are already being used in hardware engineering
Why AI infrastructure may become just as important as the models themselves
Chapters
00:00 — The AI race is shifting from intelligence to speed
00:39 — Why AI agents suddenly need faster inference
02:38 — What actually makes an AI model run faster?
05:30 — When does ultra-fast inference matter?
08:08 — Does model architecture determine inference speed?
10:05 — How do you optimize AI from model to hardware?
12:44 — How fast can AI inference actually get?
15:28 — The sequential latency problem with AI agents
16:39 — Where faster AI creates the most value
18:30 — Will premium AI inference become a competitive advantage?
20:11 — What does AI inference actually cost?
23:10 — How many tokens can AI hardware generate?
24:46 — Building private AI infrastructure
28:02 — How much faster can AI eventually become?
30:14 — Are AI agents already designing AI hardware?
32:50 — Why today's “fast” AI will eventually feel slow
33:34 — What happens next in AI infrastructure - AI attackers can now move at a speed and scale that traditional security operations were never designed to handle.In this episode of **So What About AI Agents**, Philippe Trounev sits down with **Neal Iyer, Director of Product Management for AI at Splunk Security**, to talk about what happens when autonomous AI agents become part of both the attack and defense side of cybersecurity.Neal breaks the new threat landscape into three forces: **scale, speed, and sophistication**. AI can give relatively unsophisticated attackers capabilities that previously required much more expertise, while autonomous attacks can move so quickly that a security team celebrating a five-minute detection time may already be too late.We get into:• How AI changes the economics of cyber attacks• Why cheap AI agents can dramatically increase the number and sophistication of attacks• Why traditional SOC response times may no longer be fast enough• How prompt injection creates a new attack surface for defensive AI agents• Why simply deploying more security agents isn't enough• Building a security “harness” around autonomous agents• Runtime monitoring, guardrails, observability, and human escalation• When security agents should act autonomously — and when a human needs to stay in the loop• Using business context to distinguish between compromising an intern's laptop and locking out a CEO• Why security tools need better integration for agentic operations• How organizations can start testing their agents against prompt injection and other attacks today• Why Splunk and Cisco are exploring purpose-built small language models for cybersecurity• The role of open-source, self-hosted, and multi-model AI strategies• Why AI security economics may force companies to rethink how they store and process security dataWe also get into the broader question of whether frontier AI companies can really replace specialized enterprise software — or whether impressive demos fall apart when customers need reliability, support, governance, and measurable outcomes.The bigger takeaway is that the future SOC probably isn't humans versus AI attackers.It's **agents fighting agents — with humans designing the systems, permissions, guardrails, and escalation paths that keep those agents under control.**Subscribe to **So What About AI Agents** for conversations with the people actually building and deploying AI agents in production.https://www.docsie.io
- In this episode of **So What About AI Agents**, we break down **10 AI agents and workflows we’ve used to drive traffic, generate demand, nurture leads, and automate parts of our go-to-market operation**.This episode was recorded a little while ago, so my own setup has evolved since then. Today I use **fewer, more capable agents with much more advanced workflows**, rather than trying to automate everything with a huge number of separate agents. But the core ideas in this conversation are still useful if you're trying to figure out where agents can actually create leverage in marketing and GTM.We cover agents for:• LinkedIn intent detection and outbound• Lead nurturing based on actual customer behavior• Automated newsletters• Social media content planning and distribution• Paid-ad creative and optimization• Turning videos into blogs, clips, and other content• Programmatic SEO and AI-search visibility• Building glossary and topic-cluster content at scale• Google Search Console and analytics-driven optimization• Cross-channel content repurposingWe also talk about what **doesn’t** work particularly well: generic cold email, blindly automating social media, expensive always-on AI employees, bad AI-generated content, and agents that cost more to operate than the value they create.The bigger lesson is that you probably don’t need one magical “AI employee” doing everything.You need a small number of well-defined agents with clear jobs, good data, specific triggers, and enough human oversight to keep them from doing something stupid.If you're building an AI-native marketing or GTM stack, this episode gives you a practical starting point for deciding what is actually worth automating — and what probably isn't.Subscribe to **So What About AI Agents** for conversations about how companies are actually building, deploying, and operating AI agents in the real world.https://www.docsie.io
Why Vibe Coding Fails in Production| Krishna Kumar Sharma | Ex-Amazon AI Head | Omokai EP 63
25/08/2026 | 49 minAI agents can build a demo in a day. But what happens when they touch a production system with years of technical debt, undocumented decisions, security requirements, and real customers?In this episode of So What About AI Agents, Philippe Trounev sits down with Krishna Kumar Sharma, former Head of Engineering for AI at Amazon and founder of Omokai, to talk about what agentic software development looks like outside of greenfield demos and AI hype.Krishna introduces his D3 framework — Discover, Define, Deliver — an approach to AI-assisted engineering based on the same principles used by mature software teams: understand the system, define the work, execute deliberately, and review everything.We get into:• Why greenfield AI coding demos don't represent enterprise software development• How AI-generated technical debt can compound at enormous speed• Why spawning 20, 50, or 100 agents usually isn't the answer• “Token maxing” versus ROI maxing• Using different AI models to review and challenge each other's work• Why cheaper and local models can often handle implementation after good planning• Claude, Codex, Gemini, GLM and local/edge models• Prompt caching and whether context-optimization tools actually save money• Security risks created by executives and teams vibe coding directly into production• Why human review still matters in agentic engineering• The D3 framework for AI-assisted brownfield development• Why boring, structured engineering practices become even more important with AIIn the second half, we move from software agents into the physical world.Krishna explains how Omokai is developing voice-driven command-and-control systems for robots and drones, including autonomous systems capable of operating with AI at the edge.We discuss:• Voice-controlled robots and drone swarms• Running small language models directly on robotic systems• Human-in-the-loop controls for safety-critical actions• Guardrails for autonomous machines• Robotics interfaces such as ROS2, MAVLink and PX4• Operating robots without continuous cloud connectivity• Sensor fusion, LiDAR, vision and GPS-independent navigation• Defense, security, inspection, disaster response and caregiving applications• What happens when AI agents move from software into the physical worldThe central argument of the conversation is simple:More agents aren't automatically better. More tokens aren't automatically better. The goal should be producing more value for every dollar, model call, and engineering hour you spend.Subscribe to So What About AI Agents for conversations with founders, researchers, engineers and operators actually building and deploying AI agents in the real world.https://www.docsie.ioAI Agents Won't Replace IT, They Will Redesign It - EP 62 - Shayde Christian, Cloudera
30/06/2026 | 38 minEvery CIO is asking the same question: What happens to IT when AI starts doing the work?In this episode of So What About AI Agents, Philippe Trounev sits down with Shayde Christian, SVP of Data & Analytics at Cloudera, to discuss how one of the world's largest enterprise data companies is using AI internally—not just to automate tasks, but to redesign how IT operates.Instead of focusing on AI hype, this conversation explores what actually happens inside a large enterprise when AI agents become part of daily operations.Topics include:• How Cloudera built internal AI agents for enterprise workflows• Why AI assistants and autonomous agents are fundamentally different• AI governance, testing, and production deployment• Building trustworthy enterprise AI systems• Why Cloudera reinvested AI productivity instead of laying off employees• How data teams are evolving into AI engineering teams• Measuring ROI from enterprise AI• The future role of IT departments• What CIOs and technology leaders should be preparing for todayIf you're responsible for enterprise AI, digital transformation, IT leadership, or building AI products, this episode offers practical lessons from real production deployments—not theory.GuestShayde ChristianSVP, Data & AnalyticsClouderaLinkedIn:https://www.linkedin.com/in/shaydechristian/Subscribe for weekly conversations with CTOs, CIOs, AI founders, enterprise architects, and technology leaders building production AI systems.Chapters00:00Introduction to AI Agents and Cloudera02:35The Role of AI Agents in Data Management05:46Challenges in Building AI Agents08:14AI Test Beds and Governance11:13Redesigning Roles in the Age of AI14:16The Future of AI in Business Workflows16:48AI Trust and Human Interaction19:39Agent TAM and Decision Intelligence22:28Governance and Accountability in AI25:21Concrete ROI Examples from AI Agents28:18The Future of IT and AI Integration30:50Final Thoughts and Advice for Leaders#AI #EnterpriseAI #Cloudera #CIO #ITLeadership #DataAnalytics #ArtificialIntelligence #AIAgents #Automation #digitaltransformation https://www.docsie.io
Plus de podcasts Technologies
Podcasts tendance de Technologies
Ă€ propos de So What About AI Agents
🎙 What About AI Agents is your go-to podcast for exploring the rapidly evolving world of AI agents. From automating workflows to revolutionizing industries, we break down the latest advancements, real-world applications, and emerging trends in AI.
Join us weekly as we uncover how AI agents are shaping our future, featuring expert interviews, thought-provoking insights, and stories that bridge the gap between humans and intelligent systems. Whether you're an AI enthusiast, industry professional, or simply curious about the tech shaping tomorrow, What About AI Agents has something for you.
Site web du podcastÉcoutez So What About AI Agents, Tech&Co, la quotidienne ou d'autres podcasts du monde entier - avec l'app de radio.fr

Obtenez l’app radio.fr
 gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités
Obtenez l’app radio.fr
 gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités


So What About AI Agents
Scannez le code,
Téléchargez l’app,
Écoutez.
Téléchargez l’app,
Écoutez.






































