Updated Jul 25 2026 at 12:06 AM ET

Agentic engineering with Claude Code, Cursor, Codex, and friends

Strategy, founder/researcher interviews, and industry analysis

Data for the Real World
Y Combinator 49m ago
The Advantage AI Has Over Human Mathematicians - Adam Brown
Dwarkesh Patel 5h ago
New Operating Systems for the Physical World
Y Combinator 7h ago
How Two French Engineers In New York Built The Company That Monitors The Entire Cloud
Y Combinator 8h ago
The Best One-Line Definition of an AI Model
Last Week in AI 13h ago
Opencode CEO: Blocked, 20X Growth in 6 Months, Building the Coding Agent for the World
Y Combinator 14h ago
A Field Guide to AI Freakouts
The AI Daily Brief 1d ago
Why You'll Never See Anything Enter a Black Hole - Adam Brown
Dwarkesh Patel 1d ago
Just How Good is GPT 6 Going to Be
The AI Daily Brief 1d ago
The Self Driving Company
The AI Daily Brief 2d ago
Nuclear Power Is Still Incredibly Inefficient - Adam Brown
Dwarkesh Patel 2d ago
Building a Company in Stealth | Travis Kalanick with a16z
a16z 2d ago
AI as “Weapons of Mass Destruction”
Last Week in AI 2d ago
The Artillery Officer Who Solved Einstein's Equations - Adam Brown
Dwarkesh Patel 3d ago
Is Kimi K3 Really Fable Class
The AI Daily Brief 3d ago
Why Physical AI Is the Next Frontier | Applied Intuition with a16z
a16z 3d ago
When to use Claude subagents
Peter Yang 3d ago
American Dynamism Presents
a16z 4d ago
We cut Claude Code's system prompt by 80%
Peter Yang 4d ago
“Weeks = Hundreds of Millions”: The AI Release Delay Controversy
Last Week in AI 4d ago
Planning with Claude is about finding the unknowns
Peter Yang 4d ago
Last Week in AI #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Last Week in AI 4d ago
How I Plan, Build, and Run Loops with Claude Code in 40 Minutes | Thariq Shihipar
Peter Yang 5d ago
Why Every Founder Needs a Story | Replit CEO on a16z
a16z 7d ago

Latest cs.AI / cs.LG / cs.CL preprints from arXiv

3D-Aware VLMs with Implicit and Explicit Geometries

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge th...

Wenhao Li, Xueying Jiang, Quanhao Qian et al. cs.CV 1d ago
Expanding Flow Maps

Flow-based generative models have enabled remarkable progress in fast and controllable generation across continuous and discrete state spaces, yet existing parameterizations are constrained to fixed dimensions or fixe...

Sophia Tang, Pranam Chatterjee cs.LG 1d ago
GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice,...

Vedant Shah, Onkar Susladkar, Tushar Prakash et al. cs.CV 1d ago
Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$

Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence dynamics remains poorly understood. In particular, a central unresolved question is whether BB conve...

Dawei Li, Xiaotian Jiang, Mingyi Hong math.OC 1d ago
Synthetic data generation framework for quality control automation in gravure printing

Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automated surface defect detection is critical for maintaining high-quality standards...

Korota Arsène Coulibaly, Mohamed Hamlich, Khalid Hmali et al. cs.CV 1d ago
Surprisal Theory is Tautological (without Rational Grounding)

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constrain...

Ryan Cotterell cs.CL 1d ago
Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity

Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for maintaining it. However, existing sufficien...

Hongnan Ma, Yiwei Shi, Mengyue Yang et al. cs.LG 1d ago
MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or single-turn feedback, rather than organizing an entire clinical...

Qian Wu, Xinrong Zhou, Zizhan Ma et al. cs.CL 1d ago
Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble f...

Aaron Feller, Kris Deibler, Maxim Secor cs.LG 1d ago
Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana

A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transmission patterns. Anomalies were highly structured in space and time. Ashanti a...

T. Ansah-Narh, Y. Asare Afrane cs.AI 1d ago
Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when t...

Baihui Wang, Bernard Koch cs.AI 1d ago
OpenForgeRL: Train Harness-native Agents in Any Environment

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also mak...

Xiao Yu, Baolin Peng, Ruize Xu et al. cs.AI 1d ago
Visual Contrastive Self-Distillation

On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the sel...

Yijun Liang, Yunjie Tian, Yijiang Li et al. cs.CV 1d ago
MIRROR: Learning from the Other View for Multi-Modal Reasoning

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined...

Wen Ye, Yuxiao Qu, Aviral Kumar et al. cs.AI 1d ago
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quali...

Dongjie Fu, Di Cao, Xize Cheng et al. cs.LG 1d ago
Neural solutions of coupled ghost and gluon Dyson--Schwinger equations in Landau gauge

The coupled ghost and gluon Dyson--Schwinger equations (DSEs) of four-dimensional Landau-gauge Yang--Mills (YM) theory are solved with a neural representation trained only from renormalized equation residuals. The neu...

Rodrigo Carmo Terin hep-ph 1d ago
The Boundaries of Automation: A Theory of Persistent Human Participation

The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Implicit in this pursuit is the assumption that humans remain in the l...

Fares Fourati, Hinrich Schütze, Eyke Hüllermeier et al. cs.AI 1d ago
Zero-Flow Two-Sample Tests

We propose a new approach to two-sample testing for deciding whether two sets of samples are drawn from the same distribution. The test is built on a statistical discrepancy based on the zero-flow criterion, termed ze...

Yakun Wang, Leyang Wang, Song Liu et al. cs.LG 1d ago
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one mono...

Paul Azunre cs.CL 1d ago
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) dra...

Alagappan Valliappan cs.LG 1d ago
From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs

Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing interleavings. Large language models can synthesize executable Rust tests, but their outputs often...

Kaiwen Zhang, Guanjun Liu cs.SE 1d ago
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models an...

Yueyi Liu, Chi Zhang, Sen Cui et al. cs.CV 1d ago
GS-Agent: Creating 4D Physical Worlds With Generative Simulation

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effo...

Hongxin Zhang, Chunru Lin, Junyan Li et al. cs.RO 1d ago
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified...

Linjun Li cs.AI 1d ago
🖥️ NUC-Lab · Ollama v0.30.10 + Gemma 4 / Qwen 3.6 confirmed working on RTX 5070 Ti class hardware (ASUS NUC 15 Pro, 96 GB DDR5)