Meta's 24-Hour Coding Agent
Ep 18: Meta's 24-Hour Coding Agent
Meta released Muse Code, a terminal coding agent powered by Muse Spark 1.2. It can split a large job across persistent background agents in isolated worktrees, leaving the user's working copy untouched until changes are ready for review.
The more important feature is not another benchmark score. It is the local append-only event log. Muse Code records approvals, model calls, edits, and tool use so a long-running job can recover after a crash instead of losing its state.
Meta says one kernel-optimization demonstration used more than 1,000 tool calls and ran for as long as 24 hours. That is an ambitious operating horizon, but it is still a company demonstration rather than proof of reliability on ordinary production repositories.
That distinction drives Episode 18.
Signal or Noise
- Meta Muse Code and Muse Spark 1.2, SIGNAL. Persistent agents, isolated worktrees, recoverable state, and a durable audit trail make the harness worth examining.
- OpenAI's Astra math results, SIGNAL with a condition. OpenAI released a paper and Lean certificates for ten results. Outside experts still need to check the claims.
- DeepSeek V4 Flash 0731, SIGNAL. A low-cost API with a 1 million-token context window and direct support for common agent tools.
- US frontier-model review framework, SIGNAL. The reported exemption for open models could influence how labs package and distribute future systems.
- Google DeepMind leadership reset, SIGNAL with a condition. Hassabis becomes DeepMind chair and Alphabet chief scientist while Dean leaves Google. The product evidence will be research retention and future Gemini releases.
Stack Check
- Oscar, Codex in the ChatGPT desktop app: Test whether research, planning, files, and implementation stay coherent when they share one workspace.
- Matt, Warp terminal: Run a coding agent inside it and measure whether block history, vertical tabs, and diff review reduce tool switching.
Closing takes
Oscar: Long-running agents need recoverable state and an audit trail that survives the process that created it.
Matt: The government review process is the first tax. Voluntary reviews for closed models will not stay narrow; thresholds will move, delays will grow, and open weights will become the next target under the banner of safety.
Your hosts
- Oscar Gallo: AI Engineer and entrepreneur. He lives in the intersection of engineering and businesses.
- Matt Wozniak: Serial Builder and relentless executor. He comes from the lens of what works and what doesn't.
Listen now
Sources
Like how we think about AI?
Human In the Loop is me and Matt thinking out loud. Putting that thinking to work inside a real company is the day job. If you're a founder or team trying to ship AI that survives production, let's talk.
Free 30-minute call · No pitch, just a plan · No commitment until you say go
More from the archive

Self-Hosted AI with Jackson Oaks. Why 80% of AI Bills Are Wasted.
Every fifth episode of Human In the Loop is a guest deep-dive. This week: Jackson Oaks, founder of Recursion AI and the self-hosted AI platform Courier. Jackson has…

AI Made Engineers Faster. Leadership Fell Behind
Ep 17: AI Made Engineers Faster. Leadership Fell Behind AI changed how fast a team can produce software. It did not answer the harder questions. What should you…

They Want to Ban Kimi K3
Ep 16: They Want to Ban Kimi K3 Kimi K3 became the week’s real AI story because it forced two questions into the open: how good can a downloadable model get before…