Natsuhiko Sato, Hiroshi Yoshida: An Adaptive Policy Switching Method for Episodic Cumulative Reward in Reinforcement Learning. ICMLT 2025: 135-140