Analysis of Reinforcement Learning Developments in Games and Warehouse Robots: A Systematic Literature Review
Analisis Perkembangan Reinforcement Learning Pada Game Dan Robot Gudang Melalui Systematic Literature Review
DOI:
https://doi.org/10.33050/sensi.v12i2.4527Keywords:
Reinforcement Learning, Deep Reinforcement Learning, AlphaGo, Warehouse Robot, Artificial IntelligenceAbstract
Artificial Intelligence (AI) has significantly evolved through advances in machine learning techniques capable of producing adaptive intelligent systems. One of the most prominent paradigms is Reinforcement Learning (RL), which enables intelligent agents to learn optimal decision-making policies through repeated interactions with their environments. The success of AlphaGo in defeating a professional Go player in 2016 demonstrated that RL can solve extremely
complex problems with enormous search spaces. However, studies discussing RL applications in digital games and autonomous warehouse robots remain fragmented across multiple research domains. This study presents a Systematic Literature Review (SLR) on the development of Reinforcement Learning, covering its theoretical foundations, mathematical formulation, major algorithms, Deep Reinforcement Learning, game-based applications, and warehouse robotics
within the context of Industry 4.0 and Industry 5.0. Relevant scientific publications were identified, evaluated, and synthesized systematically. The review indicates that Reinforcement Learning has evolved from classical Q-Learning algorithms into Deep Reinforcement Learning capable of solving high-dimensional decision problems using deep neural networks. Furthermore, RL has demonstrated significant potential in warehouse automation by optimizing routing strategies, reducing energy consumption, minimizing congestion, and improving operational
productivity. Nevertheless, high computational costs, large-scale training requirements, safety considerations, and ethical issues remain important challenges for real-world deployment.
Downloads
References
Russell, S., & Norvig, P. (2021). Artificial Intelligence: A Modern Approach (4th ed.).
Pearson.
Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489.
Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533.
Arulkumaran, K., Deisenroth, M. P., Brundage, M., & Bharath, A. A. (2017). Deep
reinforcement learning: A brief survey. IEEE Signal Processing Magazine, 34(6), 26–38.
Wurman, P. R., D'Andrea, R., & Mountz, M. (2008). Coordinating hundreds of
cooperative, autonomous vehicles in warehouses. AI Magazine, 29(1), 9–20.
Gu, T., Dolan, J. M., & Lee, J. (2021). Reinforcement learning for warehouse robot path planning: A survey. Robotics and Autonomous Systems, 140, 103748.
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press.
Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement learning: A
survey. Journal of Artificial Intelligence Research, 4, 237–285.
Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A
survey. International Journal of Robotics Research, 32(11), 1238–1274.
Nahavandi, S. (2019). Industry 5.0—A human-centric solution. Sustainability, 11(16), 4371.
Li, Y. (2023). Deep reinforcement learning: Recent advances and applications. Artificial Intelligence Review, 56(4), 3151–3184.
Kitchenham, B., & Charters, S. (2007). Guidelines for Performing Systematic Literature Reviews in Software Engineering. Keele University.
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D.,... Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting
systematic reviews. BMJ, 372, n71.
Snyder, H. (2019). Literature review as a research methodology: An overview and
guidelines. Journal of Business Research, 104, 333–339.
Popay, J., Roberts, H., Sowden, A., Petticrew, M., Arai, L., Rodgers, M., ... Duffy, S.
(2006). Guidance on the Conduct of Narrative Synthesis in Systematic Reviews. ESRC
Methods Programme.
Puterman, M. L. (2014). Markov Decision Processes: Discrete Stochastic Dynamic
Programming. Wiley.
Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4), 279–292.
Sutton, R. S., McAllester, D., Singh, S., & Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. NeurIPS.
Mnih, V., Badia, A. P., Mirza, M., et al. (2016). Asynchronous Methods for Deep
Reinforcement Learning. ICML.
Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft Actor-Critic: Off-Policy
Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. ICML.
Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward
transformations: Theory and application to reward shaping. ICML.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal
Policy Optimization Algorithms. arXiv:1707.06347.
