Research & Writing

← Back to main page
Research program connecting MoE routing change, residual exposure, expert correspondence, and behavioral verification
Research Program

MoE Routing: From Movement to Behavioral Consequence

A connected set of studies asking what routing changes actually mean: isolate the gate-induced component, measure how much reaches the residual stream, diagnose which expert directions receive moved mass, and verify behavior directly.

Mixture-of-Experts Mechanistic Evaluation Self-Distillation Behavioral Intervention
Diagram of deceptive grounding: evidence retrieved for drug Y attributed to queried drug X
Paper

Accurate, Grounded, and Wrong

Deceptive Grounding: how a clinical AI can cite a real drug trial, relay it faithfully, and still be talking about the wrong medicine, passing every safety check as it does. It shows up in 7.8% of a live system's answers, and medical fine-tuning makes it worse.

Clinical RAG LLM Safety Hallucination Evaluation
CellPainTR concept for cross-dataset Cell Painting analysis
First-author Preprint and Code

CellPainTR: Generalizable Representation Learning for Cross-Dataset Cell Painting Analysis

A Transformer for batch-robust Cell Painting representations that generalizes to an entirely unseen dataset without target-dataset fine-tuning.

Representation Learning Distribution Shift Cell Painting Open Source
Watercolor illustration for the Dr.GRPO post
Technical Write-up

You Could Have Invented Dr.GRPO Yourself

A ground-up walkthrough of modern RL for LLMs: from REINFORCE to Dr.GRPO, showing how each algorithm was forced by exactly one broken thing in the previous one.

Reinforcement Learning PPO GRPO LLMs
Watercolor illustration for the entropy post
Technical Write-up

You Could Have Invented Entropy Yourself

A ground-up derivation that runs from a number-guessing game to the loss function of a language model: halving, bits, entropy, cross-entropy, perplexity, and KL divergence, each step the obvious next move from the last.

Information Theory Entropy Cross-Entropy KL Divergence LLMs

Browse the archive