single-rb.php

JRM Vol.38 No.3 pp. 845-854
(2026)

Paper:

Investigation of Human Environmental Recognition During Walking Using VR and its Application to Autonomous Driving of Mobile Robots

Yamato Sato*, Haruki Ishii*, Tomokazu Takahashi*, Masato Suzuki* ORCID Icon, Kazuyo Tsuzuki** ORCID Icon, Yasushi Mae*, and Seiji Aoyagi*,† ORCID Icon

*Department of Mechanical Engineering, Faculty of Engineering Science, Kansai University
3-3-35 Yamate-cho, Suita, Osaka 564-8680, Japan

Corresponding author

**Department of Architecture, Faculty of Environmental and Urban Engineering, Kansai University
3-3-35 Yamate-cho, Suita, Osaka 564-8680, Japan

Received:
May 23, 2025
Accepted:
November 5, 2025
Published:
June 20, 2026
Keywords:
mobile robot, autonomous driving, VR, environmental survey, road detection
Abstract

Teams participating in robot competitions commonly use LiDAR-SLAM for robot navigation. However, when there are significant differences between the pre-created point cloud map of the environment and the point cloud data obtained during autonomous driving, or in open environments where acquiring point cloud data is difficult, the robot may lose its self-position and often fail in autonomous driving. Humans can reach their destinations even in unfamiliar environments by relying primarily on visual cues along with guidance information such as maps and signs. It can be said that humans rely on visual information when navigating. Similarly, if robots navigate using visual information like humans, it may become possible to achieve autonomous driving without requiring pre-acquired dense point cloud maps. In this paper, we investigated how humans walk and what they focus on while walking, both in real world and when remotely controlling a robot using VR (including cases where surrounding information other than the road was removed). The results indicated that if the recognition of the road is possible, it may be feasible to complete a course without a pre-existing map. Based on these findings, we developed a simple navigation system using road recognition using vanishing points. Its effectiveness was confirmed through driving experiments conducted on a university campus.

VR system for human-inspired navigation

VR system for human-inspired navigation

Cite this article as:
Y. Sato, H. Ishii, T. Takahashi, M. Suzuki, K. Tsuzuki, Y. Mae, and S. Aoyagi, “Investigation of Human Environmental Recognition During Walking Using VR and its Application to Autonomous Driving of Mobile Robots,” J. Robot. Mechatron., Vol.38 No.3, pp. 845-854, 2026.
Data files:

1. Introduction

In recent years, teleoperation robots and autonomous mobile robots in real-world environments where humans live have attracted attention, and research and development in this field have been actively conducted. Competitions such as the Tsukuba Challenge a and the Nakanoshima Robot Challenge b provide opportunities for demonstration experiments in urban environments to accelerate research and development. Most teams participating in these competitions adopt LiDAR-SLAM [1–3] as their navigation system. However, many teams fail to complete the course. One of the main reasons for this is that the robot loses its self-position when there are significant differences between the point cloud data obtained during driving and the pre-created point cloud map. Additionally, in open environments where the point cloud is sparse, the lack of information can also cause the robot to lose its self-position.

Humans can easily reach destinations even in unfamiliar environments. Although the oft-cited “80%” figure for visual input lacks definitive empirical proof 4, little evidence contradicts vision’s predominance among the senses. We therefore treat visual information as the primary cue for human locomotion. In this paper, we propose a navigation method based on image information instead of conventional robot navigation using point cloud data. We conducted an investigation to observe what humans focus on while walking in real world and when remotely operating a robot using VR. In VR, it is also possible to remove surrounding information other than the road and present it to the subjects. The results indicated that if the recognition of the road is possible, it may be feasible to complete a course without a pre-existing map.

Considering the above, we have developed a simple navigation system that uses vanishing points to recognize only roads and does not require a map. While previous studies using this method have reported only short distances in a straight line, our study demonstrated that, over several hundred meters, a robot can successfully continue following a road trajectory and complete navigation using only visual information. Although the applicability does not extend to environments with complex junctions such as intersections or plazas, the proposed method exhibits advantages, including low preparation cost, system simplicity, and ease of recovery, when limited to road-following scenarios. As a complementary solution to LiDAR-based SLAM, our approach serves as an effective means to expand the range of options in situations where low-cost and low-preparation deployment is required.

figure

Fig. 1. Wearing data recording camera “THINKLET.”

2. Environmental Recognition Investigation During Human Walking

First, ethics issues are explained as follows. Sections 2 and 3 of this paper describe an investigation into human walking behavior. Only fully anonymized data, excluding any personally identifiable information (such as names, addresses, or contact details), were used. According to the “Checklist for Ethical Review of Research Involving Human Subjects at Kansai University,” it was confirmed that this study does not fall under the category of “research involving human subjects” (as defined in Article 2 of the Kansai University Ethical Regulations) and therefore does not require ethical review.

To investigate environmental recognition during human walking, the subjects were asked to walk within the university campus. During the walk, a wearable device with capabilities of simultaneously recording image and audio (THINKLET, Fairy Devices) was mounted at the subjects’ eye level (Fig. 1). The subjects were asked to freely respond to two questions during their walk: “What are you consciously looking at while walking?” and “What are you consciously thinking about while walking?” (Table 1(a)). The subjects were five males, two aged 22, one aged 23, and two aged 24. Although no eye tracking devices or high-precision GPS devices were used in this experiment, reviewing the recorded images and audio after the walk allowed for a certain degree of quantitative verification of where the subjects looked and walked.

The investigation results are shown in Table 1(b) and Fig. 2. Although various responses were obtained, the most common response was that the subjects were consciously aware of the road while walking. Based on this finding, a further investigation was conducted to determine which parts of the road the subjects focused on and where they were looking (Table 2(a)). The results are presented in Table 2(b) and Fig. 3. The findings revealed that the subjects identified the boundaries of the road by paying attention to curbstones, steps, and physical obstacles such as buildings and fences. The major road features identified in the survey are illustrated in the photographs in Fig. 4.

Table 1. Pedestrian attention while walking.

figure
figure

Fig. 2. Response distribution for survey in Table 1(b).

Table 2. Road awareness while walking

figure
figure

Fig. 3. Response distribution for survey in Table 2(b).

In summary, the results of this survey revealed that humans tend to focus on the road surface, dynamic obstacles, and the surrounding environment while walking, often looking into the distance. Among these, particular attention was given to the road surface, with the subjects being especially aware of the road area and road boundary. Based on this, the hypothesis was formed that humans might be able to walk using only road area information. To further investigate human environmental recognition, an additional survey using VR was conducted in the next section.

figure

Fig. 4. Locations with the most feedback from the survey.

figure

Fig. 5. Teleoperation experimental setup.

3. Investigation of Human Environmental Recognition During Robot Teleoperation Using VR

We investigated the environmental recognition humans perform when remotely controlling a robot using VR. The robot was operated from the perspective of a camera installed on the robot. Based on the results obtained in the previous section, we also presented only the road information in VR by removing information of the surroundings.

A research and development model of an electric wheelchair (WHILL Model CR, WHILL Inc.) was used as the mobile robot equipped with sensors, a portable power supply, and a notebook PC for control (Fig. 5(a)). A stereo camera (ZED2i, Stereolabs Inc.) was used as the image sensor, and the RGB-D images obtained from it were utilized (Fig. 5(b)). The VR headset worn by the subjects (Meta Quest 2, Meta Inc.) is shown in Fig. 5(c).

An overview of the experimental method is shown in Fig. 6. The PC used to control the robot outdoors and the PC used by the subjects to view the VR images indoors are connected to the same wireless LAN available on the university campus. Through LAN communication, video and robot movement commands (for steering angles and speed) are sent and received, enabling remote control of the robot. The development of this software utilized ROS packages [c–e]. This experimental method not only ensures that the subjects can safely engage in the experiment but also provides the advantage of using robot-perspective video to investigate how humans process visual information.

The subjects wore VR headsets and conducted teleoperation experiments in two patterns: one using the standard video and the other using a video that displayed only the road area (Figs. 7(a) and (b)). After the experiment, a questionnaire was conducted asking about “whether the course could be completed using only the road area video,” and “the differences in difficulty between the two teleoperation patterns, as well as their honest impressions and observations” (Table 3(a)).

The survey results are shown in Table 3(b). All five subjects successfully completed the designated course using only the road area video, while avoiding obstacles. This demonstrates that, at a minimum, humans can remotely control a robot with just road information. However, there were gap in assumed vs. actual robot position, since landmarks were unavailable. Also, the corners were difficult to judge. It was found that the environmental information other than road was also useful for easy and secure human walking.

figure

Fig. 6. Overview of robot teleoperation experiment method.

figure

Fig. 7. VR screen images used in the teleoperation experiment.

Table 3. Environmental awareness in teleoperation.

figure

Based on the survey results in Table 3, it is believed that the visual information used by humans during remote control is categorized into levels based on their importance (Fig. 8). These levels are primarily divided into three categories, with the complexity of the information increasing as the level rises. Level 1 involves the recognition of the road area. As long as the road is visible, humans were able to control the robot remotely. Level 2 is the recognition of obstacles, which is directly related to safety during walking. In the remote control experiment, as soon as pedestrians’ feet or bicycles appeared in the video, the subjects instructed the robot to avoid them. Finally, Level 3 involves the completion of self-positioning using semantic information. Information about the surroundings, such as buildings and signs, is used to estimate the robot’s location and serves as a reference for determining self-position.

figure

Fig. 8. Levels of visual information in human walking.

In the context of autonomous driving for the robot, it is believed that by relying on Level 1 road area information for navigation, using Level 2 obstacle information to avoid hazardous objects, and utilizing Level 3 environmental information to estimate self-position and follow the correct path, it is possible to realize robot navigation that mimics human environmental recognition.

figure

Fig. 9. Road area detection program.

4. Navigation with Image Information Based on Investigation Results

As described in Sections 2 and 3, humans can walk or remotely operate robots using only road area information (Level 1). It was also found that humans tend to gaze at the distant part of the road while walking (Table 1(b)). If the robot can navigate using only the road surface region, then in road-following environments, it is possible to achieve low-preparation-cost autonomous navigation without the need for prior construction or matching of dense point cloud maps, as required in conventional LiDAR-SLAM approaches.

Since the subjects were aware of the road area and its distance (vanishing point), we created a program to calculate the vanishing point of a road (Fig. 9) 5,6,7,8,9,10,11,12. The software was implemented using the OpenCV package [f–i]. It detects the boundary lines of the road edges captured by the camera 7,8,9,10,11 (Fig. 9(b)) and identifies the location with the highest intersection density of these lines as the road’s vanishing point 7,12 (Fig. 9(c)). The robot’s navigation method is shown in Fig. 10. Software was designed to adjust the robot’s posture to ensure that the vanishing point remains on the centerline of the video frame. The actual navigation results are presented in Figs. 11 and 12.

figure

Fig. 10. Navigation method based on survey results.

figure

Fig. 11. Autonomous driving on straight and gentle curve: success.

figure

Fig. 12. Autonomous driving on L-curve: failure.

On a gently curved road like the one shown in Fig. 11, the robot successfully maintained its position in the center of the road, preventing a collision with the hedge that would have occurred if it had continued straight. However, at the L-shaped intersection in Fig. 12, the camera failed to capture the curb or step, leading to unsuccessful line detection and vanishing point calculation, resulting in a failure of autonomous driving. To address this issue, an additional algorithm was implemented, enabling the robot to rotate in place at locations where road boundary detection is difficult, as shown in Fig. 12, until it successfully detects the road boundary and recalculates the vanishing point (Fig. 13). Once the road boundary and vanishing point are detected again, the robot resumes its navigation using the method described in Fig. 10. After implementing this algorithm, further autonomous driving experiments were conducted. As a result, the robot successfully navigated the L-shaped intersection (Fig. 14) and completed a drive along the roads surrounding the university campus (Fig. 15).

figure

Fig. 13. Turning algorithm for difficult boundary detection.

figure

Fig. 14. L-curve navigation after algorithm addition.

However, the proposed method is primarily designed for road-following, and the turning behavior at corner regions is only a supplementary mechanism, implemented in a simplified manner to maintain the intended direction of travel along the road. At present, the system does not implement autonomous decision-making for selecting a travel direction at intersections. The implementation of such a function remains a future challenge.

figure

Fig. 15. Autonomous driving using vanishing point.

5. Evaluation of the Present Method

5.1. Performance of Following Road

This method is able to achieve autonomous driving without prior map creation. Conventional methods using LiDAR or GNSS require the creation of maps and route plans in advance. In the case of LiDAR, it is necessary to drive around the course beforehand to generate a point cloud map and create waypoints. In the case of GNSS, a detailed driving map must be prepared in advance by specifying the exact waypoint positions using satellite images from Google Maps, making the pre-driving preparation process time-consuming.

In our laboratory, self-localization has also been achieved using LiDAR- and GNSS-based methods, and the robot has successfully completed navigation around the campus facilities shown in Fig. 15 13,14. The respective driving results are presented in Figs. 16 and 17. On the other hand, the proposed method successfully completed a circular course on our campus with minimal preparation. This contrast does not indicate a general superiority or inferiority of the methods, but rather represents an evaluation limited to the specific perspective of preparation cost.

A previous study on autonomous driving using road vanishing points 15 successfully achieved driving on a 50-meter straight course. In contrast, the our present study demonstrated the adaptability of the proposed method to non-linear routes by successfully completing a course of approximately 350 meters that included straight sections, gentle and large curves, slopes, and sharp turns such as L-shaped intersections.

figure

Fig. 16. Autonomous driving using LiDAR.

figure

Fig. 17. Autonomous driving using GNSS.

figure

Fig. 18. Characteristics of autonomous driving results using the vanishing point.

figure

Fig. 19. Absolute lateral deviation \(x\) for each section.

figure

Fig. 20. Error rate \(\varepsilon\) for each section.

Table 4. Error analysis in each section of autonomous driving.

figure

The deviation of the robot from the road center during autonomous driving was investigated. The approximately 350-meter course shown in Fig. 18 was divided into five sections based on road shape and surrounding environment. The robot performed autonomous driving by following the vanishing point, and its deviation from the road center was measured at 10-meter intervals. The road center was set as the coordinate origin (0), with negative values representing the robot’s position to the left of the center and positive values representing its position to the right. Additionally, defining the road width as 2\(L\) and the robot’s deviation from the center as \(x\), the error rate \(\varepsilon\) was calculated using the formula \(\varepsilon=\vert x\vert/L\times 100{\%}\). The measurement results are shown in Figs. 19 and 20. Table 4 summarizes the maximum, average, and standard deviation of the error rates for each section.

On relatively narrow roads with straight sections and gentle curves, like section 1, the error rate was low, and the robot maintained its position near the road center. On the other hand, in section 2 (L-shaped road) and section 4 (large curve), indicated in yellow in Fig. 18, the error rate was higher, and the robot deviated from the road center. This deviation occurred because the robot lost sight of the road boundary. To regain the road boundary, the robot subsequently performed a turning motion.

Additionally, higher error rates were observed in section 3 and section 5. This was because, after performing turning maneuvers in section 2 and section 4, the robot temporarily moved closer to the edge of the road. In section 5, however, the robot gradually returned to the center of the road after traveling a certain distance, resulting in a decrease in the error rate.

Based on these results, it was found that the proposed method, which follows the vanishing point obtained from the road boundary, tends to exhibit significant short-term deviations at corners, but it demonstrated good driving performance on straight road sections.

5.2. Countermeasure to Cope with Intersections

While the proposed method exhibits its highest effectiveness in continuous or straight road segments as described above, the functionality for determining the turning direction at intersections with branching points has not yet been implemented. In this study, the term “map-free” refers to navigation without relying on pre-built dense metric maps such as point cloud maps, and this concept is compatible with the use of prior information necessary for directional decision-making at branching points, such as predefined instructions (e.g., “turn left at the \(N\)-th intersection”) or references to simple feature images.

To handle routes that include intersections, it is important to implement two separate modules: one for detecting the robot’s arrival at an intersection and another for determining the appropriate turning direction. In future work, we plan to integrate these modules by detecting intersection arrival based on the recognition of landmarks or characteristic features and determining the proper turning direction accordingly, thereby extending the applicability of the proposed method to more complex route environments.

5.3. Possibility of Sidewalk Compatibility

The experiments in this study were conducted on roads within the Kansai University campus. However, in public road demonstration experiments of sidewalk-driving robots, it is desirable to evaluate their driving performance on actual sidewalks c. In addition to the limited vehicular traffic, the roads used in this study are characterized by continuous road geometry formed by road shoulders, curbs, and building walls. These environmental features provide visual cues equivalent to those found on sidewalks, such as parallel boundaries formed by the road surface and curbs, walls on one or both sides, and low traffic density. However, whether the proposed method can truly operate on sidewalks cannot be determined without conducting experiments in such environments, which remains an important issue for future work.

6. Conclusions and Future Guidelines

An environmental recognition study was conducted during human walking and robot teleoperation using VR. It was found that humans tend to gaze at the ground, dynamic obstacles, the surrounding environment, and distant areas while walking. In particular, they pay significant attention to the road surface, with the road boundaries serving as important visual information. Although this observation may seem intuitive, the significance lies in having performed the investigation and confirmed these findings through actual survey data.

Robot posture control was performed using the road’s vanishing point. As a result, the robot successfully navigated along the center of the road on straight sections and roads with gentle curves. In environments like L-shaped roads, where detecting the road boundary is difficult, the robot successfully searched for the road boundary by performing a turning maneuver. Experimental evaluation of the deviation from the road center on the university campus course showed that, while a significant deviation of over 70% was observed in curved sections, the error remained below 20% on straight sections. Compared to previous studies using vanishing point information, our method successfully navigated long distances that included curves. Under the experimental conditions of this study, advantages were observed in reducing prior preparation efforts, such as point cloud map construction and waypoint setting. On the other hand, the applicability of the proposed method is limited in environments with complex topologies, such as intersections and plazas, and it is therefore positioned as a complementary option to LiDAR-SLAM.

Although the proposed system can run the course, it cannot run in environments with many dynamic obstacles (humans, bicycles, vehicles, etc.) because it does not have a dynamic obstacle avoidance function. Future work includes adding an algorithm to recognize and avoid dynamic objects and developing a self-position estimation method using information on the surrounding environment.

Acknowledgments

A part of this research was funded by the Kansai University Educational and Research Advancement Fund for the 2025 fiscal year, under the research project titled “Development of Human-Centered Living Support Technologies to Create Enriching Indoor and Outdoor Living Spaces.”

References
  1. [1] W. Hess, D. Kohler, H. Rapp, and D. Andor, “Real-time loop closure in 2D LIDAR SLAM,” 2016 IEEE Int. Conf. on Robotics and Automation (ICRA), 2016. https://doi.org/10.1109/ICRA.2016.7487258
  2. [2] K. Charalampous, I. Kostavelis, D. Chrysostomou et al., “3D maps registration and path planning for autonomous robot navigation,” arXiv preprint, arXiv:1312.2822, 2013. https://doi.org/10.48550/arXiv.1312.2822
  3. [3] S. Niijima, Y. Sasaki, and H. Mizoguchi, “Real-time autonomous navigation of an electric wheelchair in large-scale urban area with 3D map,” Advanced Robotics, Vol.33, Issue 19, pp. 1006-1018, 2019. https://doi.org/10.1080/01691864.2019.1642240
  4. [4] H. Katoh, “Origin and future of the theory that humans have obtained 80% of information input from vision,” Tsukuba University of Technology Techno Report, Vol.25, No.1, pp. 95-100, 2017 (in Japanese).
  5. [5] M. Okutomi, S. Noguchi, and K. Nakano, “Road-region detection by computing homography matrix using stereo images,” J. of the Robotics Society of Japan, Vol.18, No.8, pp. 1105-1111, 2000 (in Japanese).
  6. [6] M. Okutomi, K. Nakano, J. Maruyama, and T. Hara, “Continuous estimation of planar region for visual navigation using sequential stereo images,” IPSJ J., Vol.43, No.4, pp. 1061-1069, 2002 (in Japanese).
  7. [7] M. Kudo and H. Takahashi, “Street gutter region extraction using projective geometric features,” ITE Technical Report, Vol.41, No.12, Session ID: AIT2017-50, 2017 (in Japanese). https://doi.org/10.11485/itetr.41.12.0_21
  8. [8] K. Tanaka and M. Okutomi, “Detection of straight lines in road scenes using stereo images,” IEICE Trans. on Information and Systems, Vol.89-D, No.8, pp. 1892-1896, 2006 (in Japanese).
  9. [9] E. Adachi, T. Nabeshima, and T. Kurita, “Lane detection by Hough transform that considers posture of car,” IEICE Technical Report, Vol.105, No.615, Session ID: PRMU2005-219, pp. 103-107, 2006 (in Japanese).
  10. [10] M. Uebayashi, Y. Onishi, M. Manabe, T. Taoka, and M. Fukui, “Development of a white line recognition system for automotive camera video,” Proc. of the 69th National Convention of IPSJ, pp. 67-68, 2007 (in Japanese).
  11. [11] E. Panfilova, O. S. Shipitko, and I. Kunina, “Fast Hough transform-based road markings detection for autonomous vehicle,” Thirteenth Int. Conf. on Machine Vision, 2020. https://doi.org/10.1117/12.2587615
  12. [12] S. Ishikawa, Y. Kobayashi, T. Kaneko, A. Yamashita, and S. Ishihara, “Study on viewpoint change image generation from in-vehicle camera images,” ITE Technical Report, Vol.37, No.36, Session ID: ME2013-88, 2013. https://doi.org/10.11485/itetr.37.36.0_15
  13. [13] N. Chen, S. Suga, M. Suzuki, T. Takahashi, Y. Mae, Y. Arai, and S. Aoyagi, “Proposal for navigation system using three-dimensional maps—self-localization using a 3D map and slope detection using a 2D laser range finder and 3D map,” J. Robot. Mechatron., Vol.35, No.6, pp. 1604-1614, 2023. https://doi.org/10.20965/jrm.2023.p1604
  14. [14] J. Xue, N. Chen, T. Takahashi, M. Suzuki, Y. Mae, Y. Arai, and S. Aoyagi, “Proposal of a navigation system using a 3D maps – Development of a navigation system that integrates self-position estimation using 3D maps and route planning using 2D maps –,” Proc. of JSME Annual Conf. on Robotics and Mechatronics (Robomec2022), Session ID: 2P1-I11, 2022. https://doi.org/10.1299/jsmermd.2022.2P1-I11
  15. [15] H. Nakamura, Y. Nomura, and S. Yamamoto, “Autonomous traveling of a mobile robot with significant point images – Experiment on recognition of a traveling path –,” Proc. of the 53rd Japan Joint Automatic Control Conf., Session ID: 573, 2010. https://doi.org/10.11511/jacc.53.0.278.0
  16. [a] “Tsukuba Challenge 2024.” https://tsukubachallenge.jp/2024/ [Accessed January 3, 2025]
  17. [b] “Nakanoshima Robot Challenge 2024.” https://www.nakanoshima-rc.jp/2024/nakanoshima2024.html [Accessed January 3, 2025]
  18. [c] National Police Agency, “Guidelines for Public Road Demonstration Experiments of Sidewalk-Driving Robots.” https://www.npa.go.jp/bureau/traffic/selfdriving/roadtesting/202307robot_shiryou.pdf [Accessed October 5, 2025]
  19. [d] ROS. http://wiki.ros.org/ja [Accessed January 3, 2025]
  20. [e] Master_API. https://wiki.ros.org/ROS/Master_API [Accessed January 3, 2025]
  21. [f] Slave_API. https://wiki.ros.org/ROS/Slave_API [Accessed January 3, 2025]
  22. [g] Using the Video API. https://www.stereolabs.com/docs/video/using-video [Accessed January 3, 2025]
  23. [h] Image Filtering. https://docs.opencv.org/4.x/d4/d86/group__imgproc__filter.html#ga27c049795ce870216ddfb366086b5a04 [Accessed January 3, 2025]
  24. [i] Edge detection (SobelLaplacian, Canny). http://opencv.jp/opencv2-x-samples/edge_detection/ [Accessed January 3, 2025]

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Sep. 14, 2026