Aller au contenu

AI Odyssey

Anlie Arnaudy, Daniel Herbera and Guillaume Fournier
AI Odyssey
Dernier épisode

86 épisodes

  • AI Odyssey

    AGENTSCOPE: Why Bigger Models Do Not Fix Agent Debugging

    06/09/2026 | 19 min
    When an AI agent fails after dozens of steps, the final error rarely reveals where the problem began. AGENTSCOPE turns long execution traces into structured reasoning-action graphs, then checks them against ten neural invariants covering reasoning, control flow, and tool use.
    On the new AgentErrata benchmark, it raised exact failure-step localization from 1.32% to 31.35% with GPT-5.1 and more than doubled failure-type accuracy over a direct LLM judge. Yet the best exact localization score remains only 34.98%, and AgentErrata relies on injected, manually verified failures rather than organic production incidents.
    Inspired by the work of Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, and Mao Yang, this episode was created using Google's NotebookLM.
    Read the original paper here: https://arxiv.org/abs/2609.02371
  • AI Odyssey

    WikiSkill: The Memory Layer Agent Skills Were Missing

    29/08/2026 | 25 min
    🎧 WikiSkill: Why Agent Experience Needs a Memory Layer
    Google Research's WikiSkill separates raw execution traces, persistent knowledge, and executable skills. The authors report that this architecture improves skill evolution across five benchmarks and models, and that evolved skills can transfer between model families. Their ablation study attributes a 15-point average gain to giving the Skill Proposer access to the persistent wiki.
    For builders, this suggests that an agent's learning infrastructure can matter alongside model size: preserve the evidence behind a skill update, not only the final instructions. The study directly injects skills into prompts, does not evaluate retrieval or triggering, lacks automated wiki pruning, and excludes very long-horizon tasks.
    Inspired by the work of Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, and Tu Vu, this episode was created using Google's NotebookLM.
    Read the original paper here: https://arxiv.org/abs/2608.27454
  • AI Odyssey

    Harness Continual Learning: Your Agent Can Forget Without Changing Its Model

    22/08/2026 | 22 min
    AI agents can regress even when their foundation model never changes. The culprit may be the harness around the model: prompts, memories, tools, skills, and routing rules that evolve after every task. This episode explores Harness Continual Learning, a framework that treats this external state as the real object of adaptation. It introduces harness-level forgetting, four jointly versioned components, and a guarded proposal, evaluation, and commit loop designed to preserve reliable behavior while adding new capabilities. The paper reports gains above 10% over several baselines, but also shows that more permissive updates do not always produce a stronger final agent.
    Inspired by the work of Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, and Yang Gao, this episode was created using Google's NotebookLM.
    Read the original paper here: https://arxiv.org/pdf/2608.19013
  • AI Odyssey

    Prompting Is Dead. Loops Are the New Interface.

    06/07/2026 | 23 min
    The next frontier in AI is not better prompts. It is systems that trigger, act, observe, judge, and stop on their own. This episode explores loop engineering: the shift from manual chat with an AI to autonomous workflows that can test software, review documentation, simulate users, inspect screenshots, fix errors, and open pull requests while humans sleep.
    But autonomy has a cost. Without hard stop conditions, independent verification, maker-checker separation, and spending limits, loops can burn tokens, produce quiet technical debt, or drift into days of useless activity.
    Inspired by recent analyses from Matthew Berman, Nate Hunter, and the Prompt Engineering channel, this episode was created using Google's NotebookLM. Source note: this episode is based on multiple technical videos and developer discussions.
  • AI Odyssey

    AI Agents Are Not Agents Yet

    27/06/2026 | 22 min
    What if today’s “AI agents” are mostly automation pipelines wearing a more ambitious label?
    This episode explores Critique of Agent Model, a paper that draws a sharp line between agentic systems, which look autonomous because engineers scaffold workflows around them, and agentive systems, where goals, identity, decisions, self-regulation, and learning are internal to the system itself.
    The authors propose a Goal-Identity-Configurator (GIC) architecture as a path toward genuine machine agency, while keeping the central safety question unavoidable: greater autonomy also makes oversight significantly more difficult.
    Inspired by the work of Eric Xing, Mingkai Deng, and Jinyu Hou, this episode was created using Google’s NotebookLM.
    Read the original paper here: https://arxiv.org/abs/2606.23991
Plus de podcasts Technologies
À propos de AI Odyssey
AI Odyssey is your journey through the vast and evolving world of artificial intelligence. Powered by AI, this podcast breaks down both the foundational concepts and the cutting-edge developments in the field. Whether you're just starting to explore the role of AI in our world or you're a seasoned expert looking for deeper insights, AI Odyssey offers something for everyone. From AI ethics to machine learning intricacies, each episode is crafted to inspire curiosity and spark discussion on how artificial intelligence is shaping our future.
Site web du podcast

Écoutez AI Odyssey, Silicon Carne, un peu de picante dans un monde de Tech ! ou d'autres podcasts du monde entier - avec l'app de radio.fr

Obtenez l’app radio.fr
 gratuite

  • Ajout de radios et podcasts en favoris
  • Diffusion via Wi-Fi ou Bluetooth
  • Carplay & Android Auto compatibles
  • Et encore plus de fonctionnalités
Applications
Réseaux sociaux
v8.15.7 | © 2007-2026 radio.de GmbH
Generated: 9/12/2026 - 4:56:27 AM