Analysis of Reinforcement Learning Developments in Games and Warehouse Robots: A Systematic Literature Review

Analisis Perkembangan Reinforcement Learning Pada Game Dan Robot Gudang Melalui Systematic Literature Review

Authors

  • Aris Martono Universitas Raharja, Indonesia
  • Muhamad Iip Sahaepi Universitas Raharja, Indonesia

DOI:

https://doi.org/10.33050/sensi.v12i2.4527

Keywords:

Reinforcement Learning, Deep Reinforcement Learning, AlphaGo, Warehouse Robot, Artificial Intelligence

Abstract

Artificial Intelligence (AI) has significantly evolved through advances in machine learning techniques capable of producing adaptive intelligent systems. One of the most prominent paradigms is Reinforcement Learning (RL), which enables intelligent agents to learn optimal decision-making policies through repeated interactions with their environments. The success of AlphaGo in defeating a professional Go player in 2016 demonstrated that RL can solve extremely
complex problems with enormous search spaces. However, studies discussing RL applications in digital games and autonomous warehouse robots remain fragmented across multiple research domains. This study presents a Systematic Literature Review (SLR) on the development of Reinforcement Learning, covering its theoretical foundations, mathematical formulation, major algorithms, Deep Reinforcement Learning, game-based applications, and warehouse robotics
within the context of Industry 4.0 and Industry 5.0. Relevant scientific publications were identified, evaluated, and synthesized systematically. The review indicates that Reinforcement Learning has evolved from classical Q-Learning algorithms into Deep Reinforcement Learning capable of solving high-dimensional decision problems using deep neural networks. Furthermore, RL has demonstrated significant potential in warehouse automation by optimizing routing strategies, reducing energy consumption, minimizing congestion, and improving operational
productivity. Nevertheless, high computational costs, large-scale training requirements, safety considerations, and ethical issues remain important challenges for real-world deployment. 

Downloads

Download data is not yet available.

References

Russell, S., & Norvig, P. (2021). Artificial Intelligence: A Modern Approach (4th ed.).

Pearson.

Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.

Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489.

Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533.

Arulkumaran, K., Deisenroth, M. P., Brundage, M., & Bharath, A. A. (2017). Deep

reinforcement learning: A brief survey. IEEE Signal Processing Magazine, 34(6), 26–38.

Wurman, P. R., D'Andrea, R., & Mountz, M. (2008). Coordinating hundreds of

cooperative, autonomous vehicles in warehouses. AI Magazine, 29(1), 9–20.

Gu, T., Dolan, J. M., & Lee, J. (2021). Reinforcement learning for warehouse robot path planning: A survey. Robotics and Autonomous Systems, 140, 103748.

Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press.

Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement learning: A

survey. Journal of Artificial Intelligence Research, 4, 237–285.

Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A

survey. International Journal of Robotics Research, 32(11), 1238–1274.

Nahavandi, S. (2019). Industry 5.0—A human-centric solution. Sustainability, 11(16), 4371.

Li, Y. (2023). Deep reinforcement learning: Recent advances and applications. Artificial Intelligence Review, 56(4), 3151–3184.

Kitchenham, B., & Charters, S. (2007). Guidelines for Performing Systematic Literature Reviews in Software Engineering. Keele University.

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D.,... Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting

systematic reviews. BMJ, 372, n71.

Snyder, H. (2019). Literature review as a research methodology: An overview and

guidelines. Journal of Business Research, 104, 333–339.

Popay, J., Roberts, H., Sowden, A., Petticrew, M., Arai, L., Rodgers, M., ... Duffy, S.

(2006). Guidance on the Conduct of Narrative Synthesis in Systematic Reviews. ESRC

Methods Programme.

Puterman, M. L. (2014). Markov Decision Processes: Discrete Stochastic Dynamic

Programming. Wiley.

Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4), 279–292.

Sutton, R. S., McAllester, D., Singh, S., & Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. NeurIPS.

Mnih, V., Badia, A. P., Mirza, M., et al. (2016). Asynchronous Methods for Deep

Reinforcement Learning. ICML.

Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft Actor-Critic: Off-Policy

Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. ICML.

Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward

transformations: Theory and application to reward shaping. ICML.

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal

Policy Optimization Algorithms. arXiv:1707.06347.

Downloads

Published

2026-08-20

Most read articles by the same author(s)

1 2 > >>