Back to media
Human In the Loop · EP 18

Meta's 24-Hour Coding Agent

Podcast EpisodeAugust 11, 2026
PodcastAI AgentsCoding AgentsMeta AI
In this episode

Ep 18: Meta's 24-Hour Coding Agent

Meta released Muse Code, a terminal coding agent powered by Muse Spark 1.2. It can split a large job across persistent background agents in isolated worktrees, leaving the user's working copy untouched until changes are ready for review.

The more important feature is not another benchmark score. It is the local append-only event log. Muse Code records approvals, model calls, edits, and tool use so a long-running job can recover after a crash instead of losing its state.

Meta says one kernel-optimization demonstration used more than 1,000 tool calls and ran for as long as 24 hours. That is an ambitious operating horizon, but it is still a company demonstration rather than proof of reliability on ordinary production repositories.

That distinction drives Episode 18.

Signal or Noise

  1. Meta Muse Code and Muse Spark 1.2, SIGNAL. Persistent agents, isolated worktrees, recoverable state, and a durable audit trail make the harness worth examining.
  2. OpenAI's Astra math results, SIGNAL with a condition. OpenAI released a paper and Lean certificates for ten results. Outside experts still need to check the claims.
  3. DeepSeek V4 Flash 0731, SIGNAL. A low-cost API with a 1 million-token context window and direct support for common agent tools.
  4. US frontier-model review framework, SIGNAL. The reported exemption for open models could influence how labs package and distribute future systems.
  5. Google DeepMind leadership reset, SIGNAL with a condition. Hassabis becomes DeepMind chair and Alphabet chief scientist while Dean leaves Google. The product evidence will be research retention and future Gemini releases.

Stack Check

  • Oscar, Codex in the ChatGPT desktop app: Test whether research, planning, files, and implementation stay coherent when they share one workspace.
  • Matt, Warp terminal: Run a coding agent inside it and measure whether block history, vertical tabs, and diff review reduce tool switching.

Closing takes

Oscar: Long-running agents need recoverable state and an audit trail that survives the process that created it.

Matt: The government review process is the first tax. Voluntary reviews for closed models will not stay narrow; thresholds will move, delays will grow, and open weights will become the next target under the banner of safety.

Your hosts

  • Oscar Gallo: AI Engineer and entrepreneur. He lives in the intersection of engineering and businesses.
  • Matt Wozniak: Serial Builder and relentless executor. He comes from the lens of what works and what doesn't.

Listen now

Sources

Your move

Like how we think about AI?

Human In the Loop is me and Matt thinking out loud. Putting that thinking to work inside a real company is the day job. If you're a founder or team trying to ship AI that survives production, let's talk.

Free 30-minute call · No pitch, just a plan · No commitment until you say go

Keep going

More from the archive