← Back to Projects & Writing

Technical Write-ups

Long-form walkthroughs of papers, algorithms, and core ideas in deep learning.

Two matched MazeBench continuations after Qwen gives the same correct answer
Research Program

A Correct Answer Is Not an Action Constraint

Explicit avoidance helped Qwen choose the right move. A checked decision record caught missing commands. What this suggests for agent training and execution.

AI Agents Model Behavior MazeBench Action Validation
Qwen between an anonymous gem and a block in the MazeBench engine
Research Program

When the Right Answer Does Not Control the Next Action

Qwen named the rewarded object, then walked toward the other one. MazeBench shows where a correct explanation stops controlling behavior.

AI Agents Model Behavior MazeBench Belief Revision
MazeBench Control Center with exploration charts and model reasoning
Research Notebook

What a Local Model Does in MazeBench

What 1,021 decisions from a local 27B agent revealed about belief inertia, representation, memory, and learning from consequences.

AI Agents Reasoning MazeBench Local LLM
Watercolor illustration for the Dr.GRPO post
Technical Write-up

From REINFORCE to Dr.GRPO

A derivation of the policy-gradient methods behind modern reasoning-model training, from REINFORCE through PPO and GRPO.

Reinforcement Learning PPO GRPO LLMs
Watercolor illustration for the entropy post
Technical Write-up

Entropy from First Principles

A number-guessing game leads to bits, entropy, cross-entropy, perplexity, KL divergence, and language-model loss.

Information Theory Entropy Cross-Entropy KL Divergence LLMs
A photo of cells stylized in the manner of Picasso.
Technical Write-up

Neural Style Transfer from Scratch

An implementation of Gatys et al. (2015), with content loss, style loss, and Gram matrices derived step by step.

Style Transfer Computer Vision CNNs