AI Odyssey

Anlie Arnaudy, Daniel Herbera and Guillaume Fournier
AI Odyssey
Dernier épisode

78 épisodes

  • AI Odyssey

    AI Agents Fail the Spreadsheet Test

    25/05/2026 | 23 min
    What happens when AI agents are asked to build the spreadsheets finance teams actually use?
    WorkstreamBench, a benchmark for end-to-end financial spreadsheet work, exposes the gap between impressive demos and professional deliverables. It tests complete multi-sheet workbooks, not single formulas or table questions.
    The benchmark scores accuracy, formula quality, and formatting, because in finance a model must be auditable, readable, and easy to modify.
    Claude Web leads with 69.1 out of 100, but even the best systems degrade as tasks become more complex. Enterprise AI still has a spreadsheet reliability problem.

    Inspired by the work of Thomson Yen, Julian Poeltl, Harshith Srinivas Gear, Yilin Meng, Joshua Fan, Adam Shen, Yili Liu, Ali Bauyrzhan, Siri Du, Haoyang Liu, Daniel Guetta, and Hongseok Namkoong, this episode was created using Google's NotebookLM.

    Read the original paper here:
    https://arxiv.org/pdf/2605.22664
  • AI Odyssey

    Hermes Agent and the Rise of Agentic Operating Systems

    16/05/2026 | 15 min
    Every forty years, the way we touch a computer changes shape. The command line gave way to the mouse. The mouse gave way to the touchscreen. And now, quietly, the screen itself is starting to disappear. In this episode, we follow Hermes, an open-source agentic operating system that hit number one on OpenRouter in ninety days, processing 224 billion tokens a day. Persistent memory, self-written skills, local-first execution: Hermes is not an app you launch, it is a digital coworker that launches things for you. And while the text interface collapses into orchestration, the voice interface is collapsing into presence: Mira Murati's Thinking Machines Lab just unveiled "interaction models" that listen, watch, and speak at the same time, in 200-millisecond micro-turns. Two paradigm shifts, one direction. The OS becomes the agent. The agent becomes the conversation.
    Inspired by recent research on Agentic Operating Systems, this episode was created using Google's NotebookLM.
  • AI Odyssey

    The Agent Question Nobody Asked: When Should AI Interrupt You?

    14/05/2026 | 18 min
    Most people assume an AI agent should ask for clarification as early as possible. This paper shows that the truth is more subtle.
    For long-horizon agents — AI systems that execute many steps over time — the value of a clarification depends on what is missing : goal, input, constraint, or context. Some answers lose value almost immediately. Others remain useful much later.
    For enterprises, this is not a UX detail. It is a governance problem : when should an agent stop, ask, and avoid compounding a bad assumption?
    Inspired by the work of Anmol Gulati, Hariom Gupta, Elias Lumer, Sahil Sen, and Vamse Kumar Subbiah, this episode was created using Google's NotebookLM.
    Read the original paper here :
    https://arxiv.org/abs/2605.07937v1
  • AI Odyssey

    AI Agents Have a Coordination Problem

    10/05/2026 | 25 min
    What if multi-agent AI systems fail less because the models are weak, and more because the agents are badly coordinated? This paper treats coordination as an architectural layer : who talks to whom, who decides, how outputs are merged, and how failures are handled.
    The authors test five coordination patterns on prediction markets and find a sharp result for builders : more agents and more debate do not automatically create better systems. In this experiment, simple ensembles and sequential pipelines beat popular orchestration patterns on the cost-quality frontier.
    Inspired by the work of Maksym Nechepurenko and Pavel Shuvalov, this episode was created using Google’s NotebookLM.
    Read the original paper here :
    https://arxiv.org/pdf/2605.03310
  • AI Odyssey

    AI Agents Are Becoming Companies

    03/05/2026 | 17 min
    What if the next leap in AI agents is not a smarter worker, but a better organisation?
    This paper introduces OneManCompany, a framework that turns scattered agents, tools, skills, and runtime configurations into managed “Talents” that can be hired, reviewed, replaced, and improved over time. Its Explore-Execute-Review loop decomposes work, assigns accountability, checks outputs, and learns from failures.
    The result is striking: 84.67% success on PRDBench, beating reported baselines by 15.48 percentage points. But the catch is equally important: this organisational intelligence costs more and is still mostly validated on software tasks.
    Inspired by the work of Zhengxu Yu, Yu Fu, Zhiyuan He, Yuxuan Huang, Lee Ka Yiu, Meng Fang, Weilin Luo, and Jun Wang, this episode was created using Google’s NotebookLM. Read the original paper here: https://arxiv.org/abs/2604.22446v1
Plus de podcasts Technologies
À propos de AI Odyssey
AI Odyssey is your journey through the vast and evolving world of artificial intelligence. Powered by AI, this podcast breaks down both the foundational concepts and the cutting-edge developments in the field. Whether you're just starting to explore the role of AI in our world or you're a seasoned expert looking for deeper insights, AI Odyssey offers something for everyone. From AI ethics to machine learning intricacies, each episode is crafted to inspire curiosity and spark discussion on how artificial intelligence is shaping our future.
Site web du podcast

Écoutez AI Odyssey, Tech Café ou d'autres podcasts du monde entier - avec l'app de radio.fr

Obtenez l’app radio.fr
 gratuite

  • Ajout de radios et podcasts en favoris
  • Diffusion via Wi-Fi ou Bluetooth
  • Carplay & Android Auto compatibles
  • Et encore plus de fonctionnalités