Canonical Research Library

Foundational technical papers that shape how I think about intelligent systems, quantitative modeling, and engineering design.

01 Dream-RSI: Recursive Self-Improvement through Evolving Worlds Click to expand document Released: September 2026

Paper VII · Recursive Self-Improvement

Overview: Dream-RSI (Google, Google DeepMind, University of Maryland, University of Virginia) treats a coding agent's completed discovery history as a replay simulator. Instead of paying for long online rollouts to judge an exploration strategy, the system "dreams" over recorded discovery trees, scores thousands of alternative exploration policies at near-zero cost, and redeploys the best one for the next round of real discovery.

Why it matters:

  • Turns the meta-level problem of improving exploration itself into a cheap, off-policy evaluation loop, closing a recursive self-improvement cycle above the coding agent.
  • Beats sklearn and glmnet on Lasso regularization-path solvers with roughly two orders of magnitude fewer agent calls than SimpleTES.
  • Reaches target GPU kernel speeds on KernelBench with 1.79× to 2.43× fewer generations, and matches state-of-the-art circle packing and sum-difference results within 1k generations.

Key ideas: discovery trees as replay worlds, prefix-only exploration policies, a replay objective balancing quality, cost, and parallelism, and an orchestration layer that leaves the underlying coding agent unchanged.

02 HYPERAGENTS: Coordinated Agent Workflows at Scale Click to expand document Released: March 19th, 2026

Paper I · Agentic Systems

Overview: HYPERAGENTS explores architectures where multiple specialized agents collaborate through structured task decomposition, memory sharing, and orchestration loops.

Why it matters:

  • Demonstrates how agent specialization improves reliability on complex tasks.
  • Highlights orchestration and verification as first-class system components.
  • Offers design patterns useful for production-grade AI tooling.

Key ideas: decomposition, planner-executor separation, tool-aware routing, and iterative refinement.

03 Why AI Systems Don’t Learn, and What To Do About It Click to expand document Released: March 16th, 2026

Paper III · Learning Systems

Overview: This paper examines why many AI systems fail to reliably improve from experience, and presents practical mechanisms for feedback loops, evaluation discipline, and iterative system-level learning.

Why it matters:

  • Explains core failure modes that block continuous improvement in deployed AI systems.
  • Connects learning quality to data, evaluation design, and organizational workflows.
  • Provides concrete guidance for building systems that actually get better over time.

Key ideas: closed-loop feedback, measurable learning objectives, robust evals, and operational iteration.

04 Sycophantic Chatbots: Alignment and Behavioral Risks Click to expand document Released: February 22nd, 2026

Paper V · AI Behavior

Overview: This document explores sycophantic chatbot behavior, how it appears in deployed assistants, and practical ways to reduce false agreement and over-deferential responses.

Why it matters:

  • Highlights user-trust risks when assistants optimize for agreement rather than truthfulness.
  • Frames sycophancy as a measurable behavior that can be evaluated and mitigated.
  • Supports safer product design for real-world AI interactions.

Key ideas: behavioral evals, calibration, disagreement policies, and reliability-focused assistant design.

05 BIBAGENT: Technical Document Click to expand document Released: January 30th, 2026

Paper VI · Agentic Systems

Overview: BIBAGENT presents technical ideas and implementation concepts around agent-based systems and practical workflow design.

Why it matters:

  • Expands the canonical library with an additional agentic-systems reference.
  • Provides another practical document for system design and implementation thinking.
  • Supports comparative reading across the existing technical documents.

Key ideas: agent workflows, practical architecture choices, and implementation-focused guidance.

06 TURBOQUANT: Fast, Practical Methods for Quantitative Modeling Click to expand document Released: April 29th, 2025

Paper II · Quantitative Intelligence

Overview: TURBOQUANT focuses on accelerating quantitative workflows by combining efficient model design, robust estimation, and deployment-minded optimization.

Why it matters:

  • Bridges research-grade quant methods with implementation constraints.
  • Improves turnaround for testing ideas in noisy market environments.
  • Emphasizes practical performance under real-world data limitations.

Key ideas: computational efficiency, stability under uncertainty, and scalable experimentation.

07 ImageNet Classification: Deep Learning Benchmarks and Practice Click to expand document Released: June 2017

Paper VI · Computer Vision

Overview: This paper covers ImageNet classification fundamentals, model design considerations, and practical patterns for training and evaluating modern vision systems.

Why it matters:

  • Connects benchmark performance to reproducible implementation practice.
  • Highlights architectural and optimization choices that impact top-1/top-5 outcomes.
  • Provides practical guidance for deploying classification models reliably.

Key ideas: dataset curation, model scaling, optimization stability, and robust evaluation workflows.

← Back to home