Back to issue
skim1h 41m

🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences

Lila Sciences team (Latent Space podcast) · Latent Space

The most novel idea I heard this week: use automated wet labs as the verifier in RL with verifiable rewards, turning physical experiments into training tokens. Lila claims 10 trillion experimentally-verified reasoning tokens across bio and materials, and backs the thesis with a case study where a 2-3 person team compressed six years of CAR-T work into six months. Long, but it's a genuine primary-source look at where AI-for-science infrastructure is heading.

  • The thesis: after the internet corpus is exhausted, lab experiments become an 'infinite token generator' — verifiable rewards grounded in physical reality.
  • In vivo CAR-T case study: 2-3 people reached preclinical non-human primate data in 6 months versus a typical 6 years and $100M.
  • Cross-domain transfer works — general models trained across all sciences often beat domain-specific models sample-for-sample.
  • Iteration speed (round-over-round learning cycles) matters more than raw experiment parallelization as the scaling dimension.
Watch on YouTube

Part of Issue Nº 002: How Anthropic ships with its own agents, evals in the trenches, and the new physics of small teams