These sources detail advanced reinforcement learning frameworks designed to improve how quadruped robots navigate difficult, real-world environments. The first source introduces a single-stage teacher-student method that utilizes skeleton information and a system-response model to achieve more natural, stable movement. The second source proposes ZSL-RPPO, a zero-shot learning architecture that eliminates the need for imitation by training recurrent neural networks directly in partially observable settings. Both research papers prioritize bridging the simulation-to-reality gap, ensuring robots can handle unpredictable terrain like stairs, oily surfaces, and grass. By employing domain randomization and specialized encoders, these frameworks enhance the robustness and adaptability of robotic locomotion without requiring extensive manual tuning. Together, they represent a shift toward more efficient training paradigms that produce versatile and resilient autonomous behaviors.