Rediscovering heuristics: A litmus test for RL in networked systems
Read summary
We introduce heuristic rediscovery as a diagnostic test for RL training, asking whether compact policies can recover strong heuristic baselines. Visualizing the return landscape, optimization path, and policy correctness helps explain why training succeeds or fails. We demonstrate the methodology in adaptive bitrate streaming (ABR) and congestion control (CC).



