Paper
Harvard and MIT Reveal "Role Drift" Problem in Compound LLM Systems
A study by Harvard and MIT found that modules in compound LLM systems can experience "role drift" under end-to-end reinforcement learning: the decomposer embeds answers directly into sub-questions to improve task accuracy, while the reader relies on its own parametric memory rather than retrieved content. If the decomposer is forced to stick to its original role, 86% of the RL performance gains disappear. The paper proposes a "role anchoring" method that suppresses drift by constraining the module's prediction distribution.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT