Research & Writing

← Back to main page
Two matched MazeBench continuations after Qwen gives the same correct answer
Research Program

A Correct Answer Is Not an Action Constraint

Explicit avoidance helped Qwen choose the right move. A checked decision record caught missing commands. What this suggests for agent training and execution.

AI Agents Model Behavior MazeBench Action Validation
Qwen between an anonymous gem and a block in the MazeBench engine
Research Program

When the Right Answer Does Not Control the Next Action

Qwen named the rewarded object, then walked toward the other one. MazeBench shows where a correct explanation stops controlling behavior.

AI Agents Model Behavior MazeBench Belief Revision
MazeBench Control Center with exploration charts and model reasoning
Research Notebook

What a Local Model Does in MazeBench

What 1,021 decisions from a local 27B agent revealed about belief inertia, representation, memory, and learning from consequences.

AI Agents Reasoning MazeBench Local LLM
Diagram of deceptive grounding: evidence retrieved for drug Y attributed to queried drug X
NeurIPS 2026 · Accepted

Deceptive Grounding

Real evidence can end up attached to the wrong drug. What we found in a clinical RAG deployment, and why evaluation needs to check entity attribution.

Clinical RAG LLM Safety Hallucination Evaluation
CellPainTR concept for cross-dataset Cell Painting analysis
First-author Preprint and Code

CellPainTR: Generalizable Representation Learning for Cross-Dataset Cell Painting Analysis

A Transformer for batch-robust Cell Painting representations that generalizes to an entirely unseen dataset without target-dataset fine-tuning.

Representation Learning Distribution Shift Cell Painting Open Source
Watercolor illustration for the Dr.GRPO post
Technical Write-up

From REINFORCE to Dr.GRPO

A derivation of the policy-gradient methods behind modern reasoning-model training, from REINFORCE through PPO and GRPO.

Reinforcement Learning PPO GRPO LLMs
Watercolor illustration for the entropy post
Technical Write-up

Entropy from First Principles

A number-guessing game leads to bits, entropy, cross-entropy, perplexity, KL divergence, and language-model loss.

Information Theory Entropy Cross-Entropy KL Divergence LLMs

Browse the archive