86 épisodes
- When an AI agent fails after dozens of steps, the final error rarely reveals where the problem began. AGENTSCOPE turns long execution traces into structured reasoning-action graphs, then checks them against ten neural invariants covering reasoning, control flow, and tool use.
On the new AgentErrata benchmark, it raised exact failure-step localization from 1.32% to 31.35% with GPT-5.1 and more than doubled failure-type accuracy over a direct LLM judge. Yet the best exact localization score remains only 34.98%, and AgentErrata relies on injected, manually verified failures rather than organic production incidents.
Inspired by the work of Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, and Mao Yang, this episode was created using Google's NotebookLM.
Read the original paper here: https://arxiv.org/abs/2609.02371 - 🎧 WikiSkill: Why Agent Experience Needs a Memory Layer
Google Research's WikiSkill separates raw execution traces, persistent knowledge, and executable skills. The authors report that this architecture improves skill evolution across five benchmarks and models, and that evolved skills can transfer between model families. Their ablation study attributes a 15-point average gain to giving the Skill Proposer access to the persistent wiki.
For builders, this suggests that an agent's learning infrastructure can matter alongside model size: preserve the evidence behind a skill update, not only the final instructions. The study directly injects skills into prompts, does not evaluate retrieval or triggering, lacks automated wiki pruning, and excludes very long-horizon tasks.
Inspired by the work of Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, and Tu Vu, this episode was created using Google's NotebookLM.
Read the original paper here: https://arxiv.org/abs/2608.27454 - AI agents can regress even when their foundation model never changes. The culprit may be the harness around the model: prompts, memories, tools, skills, and routing rules that evolve after every task. This episode explores Harness Continual Learning, a framework that treats this external state as the real object of adaptation. It introduces harness-level forgetting, four jointly versioned components, and a guarded proposal, evaluation, and commit loop designed to preserve reliable behavior while adding new capabilities. The paper reports gains above 10% over several baselines, but also shows that more permissive updates do not always produce a stronger final agent.
Inspired by the work of Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, and Yang Gao, this episode was created using Google's NotebookLM.
Read the original paper here: https://arxiv.org/pdf/2608.19013 - The next frontier in AI is not better prompts. It is systems that trigger, act, observe, judge, and stop on their own. This episode explores loop engineering: the shift from manual chat with an AI to autonomous workflows that can test software, review documentation, simulate users, inspect screenshots, fix errors, and open pull requests while humans sleep.
But autonomy has a cost. Without hard stop conditions, independent verification, maker-checker separation, and spending limits, loops can burn tokens, produce quiet technical debt, or drift into days of useless activity.
Inspired by recent analyses from Matthew Berman, Nate Hunter, and the Prompt Engineering channel, this episode was created using Google's NotebookLM. Source note: this episode is based on multiple technical videos and developer discussions. - What if today’s “AI agents” are mostly automation pipelines wearing a more ambitious label?
This episode explores Critique of Agent Model, a paper that draws a sharp line between agentic systems, which look autonomous because engineers scaffold workflows around them, and agentive systems, where goals, identity, decisions, self-regulation, and learning are internal to the system itself.
The authors propose a Goal-Identity-Configurator (GIC) architecture as a path toward genuine machine agency, while keeping the central safety question unavoidable: greater autonomy also makes oversight significantly more difficult.
Inspired by the work of Eric Xing, Mingkai Deng, and Jinyu Hou, this episode was created using Google’s NotebookLM.
Read the original paper here: https://arxiv.org/abs/2606.23991
Plus de podcasts Technologies
Podcasts tendance de Technologies
À propos de AI Odyssey
AI Odyssey is your journey through the vast and evolving world of artificial intelligence. Powered by AI, this podcast breaks down both the foundational concepts and the cutting-edge developments in the field. Whether you're just starting to explore the role of AI in our world or you're a seasoned expert looking for deeper insights, AI Odyssey offers something for everyone. From AI ethics to machine learning intricacies, each episode is crafted to inspire curiosity and spark discussion on how artificial intelligence is shaping our future.
Site web du podcastÉcoutez AI Odyssey, Silicon Carne, un peu de picante dans un monde de Tech ! ou d'autres podcasts du monde entier - avec l'app de radio.fr

Obtenez l’app radio.fr gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités
Obtenez l’app radio.fr gratuite
- Ajout de radios et podcasts en favoris
- Diffusion via Wi-Fi ou Bluetooth
- Carplay & Android Auto compatibles
- Et encore plus de fonctionnalités


AI Odyssey
Scannez le code,
Téléchargez l’app,
Écoutez.
Téléchargez l’app,
Écoutez.

















![Podcast Ai Experience [en français]](https://www.radio.fr/podcast-images/175/ai-experience-en-francais.jpeg?version=f8946612c664cb12dd0835b2ab8c8163e2dd405a)



















