SELECTED PUBLICATIONS & PREPRINTS

Research

My research connects preference-based reinforcement learning, multimodal robot-data quality, vision-language-action models, and physical deployment. The common question is how robots can learn from imperfect human-provided data while preserving the behavioral distinctions that matter in the real world.

Complete citation history is available on Google Scholar. * Equal contribution.

Overview of multimodal auditing for instruction-trajectory mismatches
IEEE RA-L · 2026

Auditing Instruction–Trajectory Mismatches in Multimodal Robot Demonstrations

Simon Holk, Ryosuke Takanami, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo, Yueh-Hua Wu, and Kei Ota

We study demonstrations that execute plausible behavior but are paired with the wrong language instruction. Multimodal Probabilistic Fusion is a training-free auditing method that combines vision and proprioceptive evidence to detect and relabel these mismatches before VLA training.

RAGDP speed and accuracy results for diffusion policies
arXiv preprint · 2025

Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training

Sodtavilan Odonchimed, Tatsuya Matsushima, Simon Holk, Yusuke Iwasawa, and Yutaka Matsuo

RAGDP retrieves relevant expert actions and injects them into an intermediate denoising step, improving the speed–accuracy trade-off of pretrained diffusion policies without distillation or additional model training.

FLoRA low-rank reward-function style adaptation
ICRA · 2025

FLoRA: Sample-Efficient Preference-based RL via Low-Rank Style Adaptation of Reward Functions

Daniel Marta*, Simon Holk*, Miguel Vasco, Jens Lundell, Timon Homberger, Finn Busch, Olov Andersson, Danica Kragic, and Iolanda Leite

FLoRA adapts pretrained robot behavior to individual preferences through low-rank style parameters in the reward function, reducing the amount of new feedback needed while retaining the original task behavior.

PREDILECT language-enhanced preference learning overview
HRI · 2024

PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning

Simon Holk*, Daniel Marta*, and Iolanda Leite

PREDILECT uses zero-shot language reasoning to extract more granular information from optional natural-language explanations accompanying pairwise preferences, improving how reward models represent what a person actually wants.

POLITE preference and temporal-highlight learning architecture
ICRA · 2024

POLITE: Preferences Combined with Highlights in Reinforcement Learning

Simon Holk, Daniel Marta, and Iolanda Leite

POLITE supplements trajectory preferences with positive and negative temporal highlights so a reward model can assign credit to the behavior that actually caused the judgment. The paper was nominated for three ICRA best-paper awards.

SEQUEL semi-supervised query-synthesis framework
ICRA · 2024

SEQUEL: Semi-Supervised Preference-based RL with Query Synthesis via Latent Interpolation

Daniel Marta*, Simon Holk*, Christian Pek, and Iolanda Leite

SEQUEL learns a latent representation of trajectory pairs and synthesizes additional preference queries through interpolation, using semi-supervised learning to extract more value from limited human feedback.