Authors - Md. Istiaq Ahmed Bhuiyan, It Ee Lee, Teong Chee Chuah, Muhammad Sheraz, Gwo Chin Chung Abstract - Two-wheeled unsteady robots have special mobility benefits, but are unstable, nonlinear devices. Conventional control algorithms frequently fail to stabilize dynamic uncertainties and variations in the system. This study presents a Deep Reinforcement Learning (DRL) model based on Deep Q-Network (DQN) algorithm to balance autonomously. The proposed architecture was trained using DQN algorithm. It uses Exponential Moving Average filter to prevent high-frequency fluctuations and allows the motor output to be smooth. Simulation results demonstrate that the DQN controller successfully and robustly stabilizes the robot under mild to moderate initial pitch disturbances of up to 18°. However, boundary stress testing at an extreme 20° initial pitch revealed a critical kinematic limitation. The evaluation confirmed that while massive pitch recovery is algorithmically possible, the extreme actuator effort required to correct the chassis induces an uncontrollable divergence in the roll angle, leaving the system highly vulnerable to roll-axis instability.