AI talent & recruiting
The AI talent market is re-pricing elite researchers toward labs that can both fund aggressive compensation and sustain genuinely frontier-scale work. Moves between frontier labs are accelerating as individuals weigh total packages against the scope of problems on offer, with some researchers relocating to pursue harder scientific bets outside traditional clusters. This echoes earlier signals that raw model progress now depends less on headcount than on a narrow set of operators who can execute at the edge of what is currently feasible. At the same time, the ecosystem still lacks systematic mechanisms for identifying and placing such people, leaving recruitment more reliant on personal networks than structured scouting. The result is a bifurcated dynamic: well-resourced labs can bid aggressively for proven talent, while ambitious researchers increasingly self-select into environments that maximize intellectual upside rather than brand or location.
AI agent & harness tooling
The core friction in agent tooling is that code harnesses remain far simpler to construct than co-work equivalents, because models already excel at generating and iterating on executable logic while struggling with the messier interfaces of diagrams, spreadsheets, and unstructured collaboration. This software-native bias limits how far Codex-style systems can stretch into general knowledge work, where the output itself is rarely just runnable code. Builders are therefore turning to modular agent frameworks to replicate high-performing Claude-like setups, often leveraging community patterns around deep agent orchestration. Specific harness choices also matter at the frontier, as seen in recent state-of-the-art runs that depended on particular scaffolding choices. The pattern suggests progress will stay uneven until harness design catches up to non-code domains rather than assuming code fluency will generalize.
MoE & model efficiency research
Fine-grained mixture-of-experts designs are emerging as the practical route to reclaim efficiency where dense models waste both data and compute. Contemporary systems already borrow their routing patterns from earlier work that demonstrated sparse activation at scale, yet most still under-exploit the granularity possible. Selective pruning techniques such as REAP now let practitioners calibrate on narrow domains—coding corpora, for instance—and discard parameters that add little value there, preserving capability in the target area while shedding overhead elsewhere. This approach directly attacks the data inefficiency that currently forces ever-larger training runs. The result is a more surgical allocation of compute that aligns model capacity with actual usage patterns rather than blanket scale. Over time such methods could shift the frontier from raw parameter growth toward deliberate sparsity and domain-specific compression.
Frontier model & infra updates
Infra breakthroughs continue to outpace model scaling itself. Tri Dao’s progression from FlashAttention—now embedded across major frontier deployments—to Mamba underscores how targeted kernel and architecture work can compress inference economics far faster than raw hardware gains alone would suggest. This trajectory echoes earlier efficiency leaps that repeatedly reset the cost curve for production workloads. The result is sustained pressure on providers to either absorb lower per-token margins or accelerate release cadence to maintain differentiation, even as training tricks and learned state representations improve output coherence without proportional compute increases. Such dynamics favor teams that treat systems research as a first-class product lever rather than a downstream optimization.
Non-AI bookmarks
The through-line in these notes is the recurring friction between formal systems and the human judgment required to navigate them. Foundational advances in fields like obstetrics once met institutional dismissal before gaining traction, a reminder that novel signals often register first as eccentricity. Similar gaps appear in the distance between parsing dense corporate filings and decoding personal states, or between the nominal publicity of local statutes and their practical accessibility. Requests for obscure reading material and the deliberate retention of professional boundaries reflect the same underlying discipline: cultivating inputs and relationships that resist easy categorization. In each case the work lies less in acquiring more data than in sustaining the attention needed to interpret what existing structures obscure.
Research questions
- What is the actual retention elasticity for elite researchers when frontier labs offer 2-3× comp jumps versus equity in Series B/C startups, and which labs are winning the last 12 months of moves?
- How much inference-cost reduction is achievable by combining fine-grained MoE routing with aggressive pruning before the quality cliff appears, and which open-weight releases are already demonstrating this curve?
- Do forward-deployed engineers embedded at key customers create durable product moats, or do they simply accelerate bespoke implementations that later get productized by pure agent-harness startups?
- Which post-training techniques (RL self-play, synthetic data loops, decision-trace harvesting) are showing the steepest slope on agent reliability benchmarks, and are any labs open-sourcing the harnesses that produced those gains?
- How are export-control regimes affecting the cross-border flow of both model weights and the researchers who know how to train them—specifically, which Chinese labs are still able to recruit Western talent and at what comp premium?
Momentum
- Agent harnesses & infrastructure — BUILDING (day 3 running): threads on code harnesses, co-work agents, and decision-trace tooling recur from 06-15 through 06-21 with increasing specificity.
- AI talent & recruiting — NEW: first explicit focus on high-profile moves, comp jumps, and elite-researcher hiring appears today; prior days touched labs/geopolitics but not compensation dynamics.
- MoE & model efficiency research — NEW: fine-grained MoE, pruning, and data/compute efficiency surface for the first time; baseline just starting.
- Frontier model & infra updates — FADING: new model releases, training tricks, and infra cost curves were prominent 06-14 to 06-15 (Anthropic Fable/Mythos, open vs closed strategy) but absent or thinned since 06-16.
- Export controls / Anthropic Fable guardrails — FADING: dominant 06-14 through 06-17 (export ban, jailbreaks, regulatory fallout) but no direct mentions in 06-20 or 06-21.
AI-synthesized from your bookmarks; quotes are paraphrased and linked to source. Sanity-check any figure before citing it elsewhere.