Portrait of Cédric Caruzzo

Cédric Caruzzo

AI Research Scientist at Lunit

Reliable foundation models under distribution shift

Representation learning · Post-training · Evidence-grounded biomedical AI

Selected research Download CV

Research Profile

I study how foundation models behave when their data, tasks, and evidence change. My research spans representation learning, post-training, retrieval, and evaluation, with a focus on reliable biomedical AI.

At Lunit, I work on self-supervised vision models for whole-slide pathology and post-training and evaluation of large medical language models. My recent first-author work introduced deceptive grounding, a clinical RAG failure in which a model can faithfully cite real evidence while assigning it to the wrong medical entity.

At KAIST, I developed CellPainTR, a Transformer for cross-dataset Cell Painting analysis that generalizes to unseen datasets without target-dataset fine-tuning. Earlier, at Institut Pasteur, I built machine-learning systems for phenotypic drug discovery and infectious-disease research.

I am particularly interested in adaptive and multimodal learning, reliable post-training, and medical AI systems whose evidence and intermediate reasoning can be independently verified.

Publications and Research Outputs

Workshop paper Under review
Which Expert Directions Receive Moved Gate Mass? An Exact Correspondence Reference for MoE Routing Changes
Cédric Caruzzo, Donggeun Yoo, Tae Soo Kim
Workshop paper, under review, 2026

Defines an exact conditional correspondence reference for moved gate mass, factors it into spectral capacity and directional utilization, and shows through a preregistered reversal why geometry is diagnostic rather than a checkpoint-independent behavioral allocator.

arXiv 2026
Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation
Cédric Caruzzo, Donggeun Yoo, Tae Soo Kim
arXiv preprint arXiv:2608.15787, 2026

Introduces an exact routing-content decomposition and traces same-weight MoE routing mismatch from block output to residual exposure and causal output effects across seven checkpoints and two domains.

arXiv 2026 Under review
Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
Cédric Caruzzo, Donggeun Yoo, Tae Soo Kim
arXiv preprint arXiv:2607.09349, 2026

Introduces an entity-attribution failure missed by standard RAG evaluation, characterizes it in controlled and production settings, and proposes entity-attribution verification.

arXiv 2025
CellPainTR: Generalizable Representation Learning for Cross-Dataset Cell Painting Analysis
Cédric Caruzzo, Jong Chul Ye
arXiv preprint arXiv:2509.06986, 2025

Learns batch-robust cellular representations across heterogeneous studies and generalizes to an unseen dataset without target-dataset fine-tuning.

SLAS Europe 2024 Poster
Cellular Phenotypic Profiling: Combining Live Imaging and Cell Painting Techniques
Soonju Park, Cédric Caruzzo, Nakyung Lee,
Poster presented at SLAS Europe 2024, Barcelona, Spain

Combines longitudinal live-cell imaging with endpoint Cell Painting for richer phenotypic profiling in high-content drug discovery.

Scientific Reports 2021
Wolbachia detection in Aedes aegypti using MALDI-TOF MS coupled to artificial intelligence
Antsa Rakotonirina, Cédric Caruzzo, Valentine Ballan,
Scientific Reports 11, 21355, 2021

Uses machine learning with MALDI-TOF mass spectrometry to detect Wolbachia infection in Aedes aegypti mosquitoes for vector-control research.

Selected Experience

Lunit, Seoul
Mar 2025 - Present
AI Research Scientist
  • Own post-training and evaluation pipelines for large medical language models, including benchmark design, retrieval, and knowledge-grounded reasoning.
  • Develop self-supervised Vision Transformers for whole-slide pathology and distributed inference systems for clinical-scale image collections.
  • Lead research on evaluation blind spots in medical RAG, including the first-author Deceptive Grounding study.
KAIST, Kim Jaechul Graduate School of AI
Feb 2023 - Feb 2025
M.S. Researcher, advisor: Prof. Jong Chul Ye

Full research experience

Selected Research

Deceptive Grounding failure structure
First-author preprint · 2026

Deceptive Grounding

Clinical RAG can relay real evidence, cite a genuine source, and still attribute that evidence to the wrong medical entity. Standard hallucination and faithfulness checks miss the failure.

Clinical RAG Entity attribution Evaluation blind spot
MoE research program connecting routing change, residual exposure, expert correspondence, and behavioral verification
Research program · 2026

MoE Routing: From Movement to Behavioral Consequence

A connected research program on what routing changes actually mean: isolate the gate-induced component, quantify residual exposure, diagnose which expert-output directions receive moved mass, and establish consequence through behavioral intervention.

Residual exposure Expert correspondence Behavioral verification
CellPainTR cross-dataset representation-learning concept
First-author preprint and code · 2025

CellPainTR

A Transformer for Cell Painting representations that preserves biological signal across batch and feature shifts, including generalization to an entirely unseen dataset without fine-tuning.

Cross-dataset OOD No target fine-tuning Open source

Selected Technical Writing

All research and writing