Aller au contenu

AI Odyssey

Anlie Arnaudy, Daniel Herbera and Guillaume Fournier
AI Odyssey
Dernier épisode

90 épisodes

  • AI Odyssey

    Personal AI Agents That Act for You: Hermes, OpenClaw, Grok Bot and Muse

    27/09/2026 | 18 min
    🎧 Personal AI Agents That Act for You: Hermes, OpenClaw, Grok Bot and Muse
    AI agents are moving beyond answering questions. You can give them a task, let them work across apps, and step in when needed. This episode looks at four options for individuals. Hermes Agent and OpenClaw offer control on your own devices, with setup required. Grok Bot works in a cloud computer for eligible subscribers. Meta’s Muse is built for everyday tasks: it can use a browser and connected apps, keep working after you close the app, and asks for approval before sending an email or making a purchase, according to Meta. Muse is rolling out in the US.
    We compare access, practical uses, and permissions. We also cover the unconfirmed OpenAI DevDay rumor without treating it as a launch. This AI Odyssey episode was created with Google’s NotebookLM from many sources, including official product information and reporting.
  • AI Odyssey

    Jev: What’s Behind the Buzz?

    20/09/2026 | 16 min
    🎧 Jev: What’s Behind the Buzz?

    Jev is getting attention with a simple promise: AI that makes a choice instead of writing an answer. But what does that actually mean?

    In this episode, we unpack TypeSafe AI’s new model: what it is, how software can use it to sort requests or choose a next step, and why its claims of faster, cheaper decisions are attracting interest.

    We also look at what Jev does not solve. A neatly formatted answer can still be wrong, and the company’s performance claims need independent testing. What is useful here, and what still needs proving?

    Inspired by the work of Diogo Almeida and the TypeSafe AI team, this episode was created using Google's NotebookLM.

    Read the original announcement here: https://typesafe.ai/blog/introducing-system-one-models-and-jev
  • AI Odyssey

    Grounding Agent Memory: When AI Must Check What It Remembers

    14/09/2026 | 25 min
    🎧 Grounding Agent Memory: When AI Must Check What It Remembers
    An AI assistant that remembers yesterday can repeat yesterday’s mistakes. This episode explores research from Microsoft on checking an agent’s memories against its working environment before saving them for future tasks.
    A separate curator inspects databases or documents through read-only tools, then corrects, narrows or discards uncertain memories. In one database benchmark, success reached 73%, compared with 70% for memory alone and 39% without memory. The question is whether better verification justifies its extra background work: reported task-agent savings exclude curation costs.
    Inspired by the work of Susheel Suresh, Hazel Mak, Sahil Bhatnagar, Chhaya Methani and Alejandro Gutierrez Munoz, this episode was created using Google's NotebookLM.
    Read the original paper here: https://arxiv.org/abs/2609.11060v1
  • AI Odyssey

    Recursive Self-Improvement: AI Must Learn to Improve Its Own Learning

    13/09/2026 | 19 min
    🎧 Recursive Self-Improvement: AI Must Learn to Improve Its Own Learning

    An AI that fixes one answer has not necessarily learned anything for tomorrow. Recursive self-improvement asks for something harder: changes that persist across tasks and reshape how the system makes its next improvements.

    We explore a new research roadmap that separates five levels of autonomy, from executing prescribed updates to revising the mechanisms of improvement itself. The distinction matters for anyone deciding how much control to give an agent over its tools, training, and evaluation.

    The paper surveys emerging systems and preliminary industry evidence. It offers a framework for judging progress, not proof that fully autonomous recursive improvement has arrived.

    Inspired by the work of Yi Duan and colleagues, this episode was created using Google's NotebookLM.

    Read the original paper here: https://www.alphaxiv.org/abs/2609.11873
  • AI Odyssey

    AGENTSCOPE: Why Bigger Models Do Not Fix Agent Debugging

    06/09/2026 | 19 min
    When an AI agent fails after dozens of steps, the final error rarely reveals where the problem began. AGENTSCOPE turns long execution traces into structured reasoning-action graphs, then checks them against ten neural invariants covering reasoning, control flow, and tool use.
    On the new AgentErrata benchmark, it raised exact failure-step localization from 1.32% to 31.35% with GPT-5.1 and more than doubled failure-type accuracy over a direct LLM judge. Yet the best exact localization score remains only 34.98%, and AgentErrata relies on injected, manually verified failures rather than organic production incidents.
    Inspired by the work of Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, and Mao Yang, this episode was created using Google's NotebookLM.
    Read the original paper here: https://arxiv.org/abs/2609.02371
Plus de podcasts Technologies
À propos de AI Odyssey
AI Odyssey is your journey through the vast and evolving world of artificial intelligence. Powered by AI, this podcast breaks down both the foundational concepts and the cutting-edge developments in the field. Whether you're just starting to explore the role of AI in our world or you're a seasoned expert looking for deeper insights, AI Odyssey offers something for everyone. From AI ethics to machine learning intricacies, each episode is crafted to inspire curiosity and spark discussion on how artificial intelligence is shaping our future.
Site web du podcast

Écoutez AI Odyssey, Tech&Co, la quotidienne ou d'autres podcasts du monde entier - avec l'app de radio.fr

Obtenez l’app radio.fr
 gratuite

  • Ajout de radios et podcasts en favoris
  • Diffusion via Wi-Fi ou Bluetooth
  • Carplay & Android Auto compatibles
  • Et encore plus de fonctionnalités
Applications
Réseaux sociaux
v8.18.0 | © 2007-2026 radio.de GmbH
Generated: 9/28/2026 - 4:18:49 PM