Paper:
Remote Support and Education System for Trauma Treatment at Disaster Sites
Daichi Aoki*, Soichi Murakami**
, Chisaho Miura*, Satoshi Kanai*
, Takashige Abe***, Taku Senoo*
, Toshiaki Shichinohe**
, and Atsushi Konno*

*Graduate School of Information Science and Technology, Hokkaido University
Kita 14, Nishi 9, Kita-ku, Sapporo, Hokkaido 060-0814, Japan
**Hokkaido University Hospital
Kita 14, Nishi 5, Kita-ku, Sapporo, Hokkaido 060-8648, Japan
***Faculty of Medicine, Hokkaido University
Kita 15, Nishi 7, Kita-ku, Sapporo, Hokkaido 060-8638, Japan
This study proposes a dual-purpose system for remote medical support during disasters and trauma treatment education, aiming to enhance procedural guidance and training effectiveness in resource-limited and high-risk environments. The system consists of a remote support system for trauma treatment and a medical education system for trauma training. The remote support system enables trauma experts at distant locations to interact with a 3D digital twin of disaster victims and their surroundings within a virtual reality environment, thereby providing procedural guidance to on-site medical personnel. The system uses widely available devices such as smartphones for 3D scanning and head-mounted displays for immersive visualization. This design enables rapid deployment at disaster sites without requiring specialized equipment or complex setup procedures. The trauma education system records expert surgical treatments using motion-capture technology and reconstructs them as interactive 3D avatars, allowing trainees to observe and learn techniques from multiple perspectives. The remote support system was evaluated through fingertip-based interaction tasks in a simulated disaster scenario, where the alignment accuracy was assessed using augmented reality overlays, resulting in a measured average error of 27.0 mm. Similarly, the trauma education system was evaluated for positional accuracy in instrument handling tasks. These results confirm the feasibility and practicality of the proposed system and demonstrate its potential to improve both emergency medical response and surgical education.
Overview of the remote support system
1. Introduction
Natural disasters such as earthquakes and typhoons, terrorism, large-scale incidents, and accidents occur infrequently but result in a sudden surge in casualties. In such situations, the demand for medical care often exceeds the available supply, requiring a limited number of medical personnel to manage a large number of patients 1. Furthermore, disaster-site medical personnel may be required to provide treatment outside their areas of expertise, making it challenging to ensure appropriate medical care 2. For example, during the 2010 Haiti earthquake, many surgeons were compelled to perform surgical treatments beyond their specialization, particularly limb amputation and other surgical interventions. Subsequent surveys have revealed that many of these treatments were performed inadequately 3,4. These findings highlight the need for a support system that enables non-specialist medical personnel to deliver appropriate treatment during disaster scenarios, thereby improving the quality of emergency medical care.
Remote support systems are widely utilized in various fields, including medical care, industry, and education. Materna et al. developed a teleoperation interface for a semi-autonomous robot to assist the elderly 5. This system can be learned in a short time and is easy to operate even for non-specialists. Oliveira et al. developed a system that allows remote technicians to learn practical tasks with expert guidance through voice recognition and video guidance 6.
In addition to remote support during disasters, acquiring trauma treatment skills is essential for improving surgeons’ technical proficiency in standard medical education. According to a study by Franko et al., the number of war-related cases decreased between 2009 and 2018, with a particularly notable 50% reduction in upper limb amputations 7. Given this trend, the necessity of acquiring relevant skills to prepare for disasters and accidents has been emphasized. In Japan, young trauma surgeons can enhance their skills through programs such as the Japan Advanced Trauma Evaluation and Care (JATEC) course, organized by the Japan Trauma Care and Research Organization a. The JATEC course provides opportunities to acquire essential trauma-care knowledge and emergency treatment skills through e-learning and simulated clinical training. Furthermore, hands-on trauma training has traditionally been conducted using cadavers 8 and animal organs. The Advanced Surgical Skills for Exposure in Trauma (ASSET) course, hosted by the Japan Surgical Society, allows participants to practice surgical treatments for penetrating injuries and gunshot wounds using cadavers b. However, the availability of cadavers is restricted by ethical and legal regulations, and the number of training locations and participants is limited, making it difficult to ensure sufficient training opportunities. Animal organs differ significantly from human organs in terms of mechanical properties and dimensions, limiting their effectiveness in trauma treatment training.
Virtual reality (VR)-based training is a relatively accessible method for acquiring procedural skills. For example, Harrington et al. developed a VR simulator to train decision-making in managing patients with blunt abdominal trauma, where trainees used an Oculus Gear VR headset to review patient data in a virtual environment and select appropriate treatment options 9. This immersive experience allows trainees to engage in clinical scenarios in risk-free settings. Furthermore, the effectiveness of VR training has been demonstrated through comparative studies. Huri et al. evaluated the training outcomes of VR-based simulators versus cadaver-based education for orthopedic trainees 10. Their results showed that the VR-trained group exhibited shorter treatment times and caused less tissue damage, suggesting that VR-based training can be as effective as, or even superior to, conventional cadaveric methods.
To further enhance the realism and effectiveness of such training systems, motion-capture technologies have been integrated into medical simulation platforms. OptiTrack is a widely used motion capture system that has been adopted in various medical research contexts to capture precise human and tool kinematics. For example, Ebina et al. developed a motion measurement system for laparoscopic surgery training using OptiTrack to evaluate the skill levels of surgeons based on instrument trajectories 11. Beyond such skill-evaluation systems, Gasques et al. proposed ARTEMIS, a collaborative mixed-reality system that combines VR, augmented reality (AR), and motion tracking to enable immersive surgical telementoring 12. Their evaluation demonstrated that, even in the presence of minor spatial misalignments between virtual guidance and physical tasks, effective communication was maintained through a combination of verbal instructions and hand gestures. This finding supports the integration of motion capture and mixed reality technologies in remote medical guidance systems, especially in high-pressure environments, such as trauma care.
The originality and contributions of this study can be summarized as follows. First, a VR/AR-based remote support system is proposed for trauma treatment specifically at disaster sites, where trauma specialists provide remote guidance. Previous studies have primarily reported remote support in stable environments, such as operating rooms 12. In the proposed system of this study, only LiDAR-equipped smartphones or tablets and commercially available 3D scanning software are assumed to be used at disaster sites. Since no dedicated or specialized equipment is required to acquire information about both victims and surrounding environments, the proposed approach provides a practical solution for real disaster scenarios. Moreover, the system can be operated by on-site responders without specialized expertise or advanced technical training.
Second, the proposed VR/AR-based remote support system introduces a quantitative evaluation of the positional accuracy in remote procedural guidance. Prior studies generally relied on visual or subjective assessments without reporting numerical accuracy 12. In this study, photogrammetry was employed to measure the error between the instructed position displayed by the AR avatar and the actual treatment location, which was marked with cross-shaped markers. This methodology enables an objective and quantitative assessment of the alignment accuracy in remote guidance.
This paper presents the design, implementation, and evaluation of two systems: (1) a VR-based remote support system for trauma treatment in disaster scenarios and (2) a motion capture-driven trauma education system. Section 2 describes the architecture of the remote support system, followed by an experimental evaluation in Section 3. Section 4 presents the medical education system and its validation. Section 5 discusses the technical contributions, limitations, and prospects of the proposed approach. Finally, Section 6 concludes the paper.
2. Remote Support System for Trauma Treatment
2.1. System Overview
An overview of the proposed remote trauma support system is illustrated in Fig. 1. The system is composed of three main components: disaster area, remote assistance Metaverse environment, and remote medical center. These three locations are connected via a workflow that enables trauma experts to perform procedures in VR and deliver visual guidance back to the disaster site via AR.

Fig. 1. Overview of the remote support system.
At the disaster site, medical personnel use a LiDAR-equipped smartphone to scan the injured victim and the surrounding debris. This process generates a three-dimensional digital twin of the environment (1), which is transmitted to the cloud-based remote assistance Metaverse. In this virtual space, remote trauma experts visualize the reconstructed disaster scene through a VR interface (2). Within the remote medical center, experts perform simulated trauma treatments, such as limb amputation or intravenous infusion, based on their expertise (3). These physical motions are captured using a full-body motion capture system and finger-tracking gloves and then reflected in a digital avatar within the Metaverse (4). The avatar’s actions are subsequently sent back to the disaster area, where they are displayed through an AR interface on mobile devices such as smartphones or tablets (6). This enables on-site personnel to visually follow the expert’s procedural demonstration in spatial alignment with the actual patient.
In this study, the entire system flow, excluding the simulated tissue deformation of the patient model, was implemented and evaluated. Steps (5) and (6) involving dynamic updates to the virtual patient model (e.g., deformation due to surgical manipulation) have not yet been realized. However, future extensions of the system will include physics-based simulation of treatment effects, such as soft-tissue deformation or bleeding, enabling the AR visualization to reflect not only the expert’s motions but also the procedural outcome on the patient in real time.
2.2. 3D Scanning and Data Processing
To construct a digital twin of the disaster site, a LiDAR-equipped Apple iPhone 16 Pro (hereinafter, iPhone) was used for 3D scanning. The scanning process employed Polycam c, a commercially available 3D scanning application that utilizes the iPhone’s LiDAR sensor and photogrammetry algorithms to generate high-resolution 3D models of objects and environments. A comparative study reported that Polycam achieves a global accuracy with a standard deviation of 7 cm, with 77% of points within \(\pm\)5 cm. A local mean error of 2.0 cm demonstrates sufficient precision for practical applications in field environments 13.
The choice of smartphone-based scanning is driven by practical considerations for disaster response scenarios. Dedicated 3D scanners are often large and expensive, and require specialized operations, making their deployment in the field impractical. In contrast, smartphones are lightweight, portable, and widely available, enabling emergency responders to acquire spatial data rapidly with minimal setup effort.
Kako et al. 14 analyzed 24 rescue operations conducted by highly skilled rescue teams during the 2016 Kumamoto earthquake. In four of these cases, “Confined Space Medicine,” including treatments such as intravenous therapy performed inside collapsed buildings, was reported. In these instances, after confirming the victim’s condition, an average of 39 minutes (minimum 20 minutes, maximum 60 minutes) elapsed from the request for medical intervention to the start of collaboration at the site.
The proposed system assumes that a 3D model of the victim and surrounding environment is generated at the stage of confirming the victim’s condition. Thus, the interval between the request for medical intervention and the arrival of medical personnel at the site (20 minutes in the shortest case during the Kumamoto earthquake) can be used for LiDAR scanning and model generation. In this experiment, three scans were performed on a simulated 1.5 m \(\times\) 2 m disaster environment, including a mannequin and debris. The average time required for scanning to generate a 3D model was 116 seconds, which was well within the above-mentioned 20-minute timeframe.
The 3D scanning process involves capturing the victim and surrounding debris on-site, generating a model in “.glb” format. Polycam generates surface mesh models either from point clouds acquired by LiDAR or from multiple photographs captured from different viewpoints. The data are then transferred from the iPhone to a PC for further processing. The game engine Unity 2022.3.13f1 d was used to reconstruct the virtual disaster environment and enable interactive visualization. Unity provides a flexible platform for integrating 3D models, motion capture data, and VR device interfaces, making it suitable for the development of real-time, immersive support systems. Since Unity does not natively support “.glb” files, the glTFast plugin was incorporated to facilitate the seamless import and rendering of the scanned models.
Although the current workflow requires manual data transfer and model integration, future improvements will focus on automating these processes to further streamline system deployment in real disaster scenarios.
2.3. Remote Expert’s VR Treatment and Motion Capture
A remote trauma expert operates within a virtual reality (VR) environment to interact with the digital twin of the disaster site and perform simulated treatment procedures. In this study, Meta Quest 3 e was employed as the head-mounted display (HMD) to provide immersive visualization of the reconstructed disaster environment. The VR system was developed using Unity 2022.3.13f1, incorporating the Meta XR Core package for device integration and the OVRCameraRig prefab for stereoscopic rendering and spatial interactions.
A wireless connection between the HMD and the workstation was established via Air Link, allowing untethered movement during the procedure. Furthermore, OpenXR was used as the runtime interface to maintain compatibility across VR platforms and ensure stable system performance.
To capture the expert’s body movements accurately, an optical motion capture system comprising 12 OptiTrack PrimeX41 f cameras was employed. Fifty infrared reflective markers were attached to a motion capture suit worn by the expert, following a standard full-body skeleton configuration. The spatial positions of these markers were tracked in real time using Motive 3.0.3 software, and the resulting motion data were streamed into Unity via the official OptiTrack Unity Plugin. This setup enabled seamless synchronization between the expert’s physical movements and their digital avatar within the VR environment.
MANUS METAGLOVES PRO g was used for detailed finger articulation. These gloves use absolute-position fingertip sensors based on quantum-tracking technology to deliver drift-free, millimeter-accurate finger motion capture. A reference sensor at the back of the hand enables accurate pose estimation through sensor fusion. Combined with full-body motion capture, this setup enables the expert to perform realistic treatment procedures within the VR space.
An overview of the VR treatment and motion capture integration of the expert is shown in Fig. 2. This integrated system allows the remote expert to intuitively demonstrate trauma treatment procedures, with their body and hand movements accurately reflected by a digital avatar in the VR environment. The captured interactions form the basis for the procedural guidance provided to on-site medical personnel, which will be visualized and evaluated in subsequent sections.

Fig. 2. Treatment environment for remote expert.
2.4. Avatar Modeling and Virtual Treatment

Fig. 3. Flow of the avatar creation process.
To represent the remote expert’s body movements in a virtual environment accurately, a digital avatar was created using a combination of body measurement data, skeletal tracking, and mesh modeling. This avatar served as the medium through which the expert demonstrated trauma treatment procedures within the system.
The avatar was generated using DhaibaWorks 15, a digital human modeling platform developed by the National Institute of Advanced Industrial Science and Technology (AIST). Body feature points obtained from the motion capture system were used to fit a template model, resulting in a personalized skin mesh that reflects the expert’s actual body proportions. This ensured that the link lengths and overall dimensions of the avatar accurately matched those of the operator.
Figure 3 shows the data flow of the avatar-creation process. Skeletal data were exported from Motive 3.0.3 in “FBX” format, capturing the joint hierarchy and motion information. The exported skeletal model and skin mesh generated by DhaibaWorks were then imported into Blender h, which is a 3D computer graphics software package. In Blender, rigging and bone alignment adjustments were performed to ensure consistency between the captured motion data and avatar movement. The finalized avatar model was subsequently imported into Unity, where it was linked to real-time motion data streamed from the motion capture system.
Within the virtual environment, the expert’s avatar was visualized from multiple perspectives to support different procedural needs. A first-person view (FPV) was provided to the expert through Meta Quest 3, which enabled immersive and intuitive interaction with the virtual disaster environment. Additionally, a third-person view (TPV) was implemented within Unity for external observations and system verification.
For documentation and analysis, the system’s output was recorded through three channels:
-
(a)
External video footage recorded using a video camera observing the expert’s physical movements during treatments.
-
(b)
TPV captured within the Unity environment.
-
(c)
FPV rendered in Unity and displayed via the HMD.
An example of these visual perspectives is presented in Fig. 4.

Fig. 4. Experiment on remote support for trauma treatment.
2.5. AR-Based Result Display

Fig. 5. AR-based visualization of an expert’s trauma treatment. An avatar reproducing the expert’s trauma treatment actions is superimposed onto a video of a simulated disaster site created with rubble and a mannequin.
An AR visualization system was developed to deliver procedural guidance to on-site medical personnel intuitively. By superimposing the expert’s procedural demonstration onto a real-world environment, the system enabled non-specialist medical personnel to follow trauma treatments with spatial and contextual clarity.
AR visualization was implemented using Unity, integrating AR Foundation for cross-platform AR development and Immersal SDK for spatial mapping and localization.
First, the simulated disaster environment (mannequin and debris) was scanned using Polycam to generate a high-resolution 3D model. The data were uploaded to the Immersal i server to create a feature point map and the corresponding 3D spatial environment. Immersal’s visual positioning system enables the localization of AR content by matching real-world camera views to a pre-scanned feature point map.
Within Unity, the expert’s avatar, animated based on recorded procedural demonstrations, was integrated into the AR scene. This allowed disaster-site personnel to view the expert’s trauma treatment in situ, overlaid on the physical environment, using a smartphone or tablet. For deployment, the AR application was built using Xcode and installed on an iOS device.
An example of the AR visualization is shown in Fig. 5. To enhance usability, basic playback controls, such as play, pause, and seek bar functions, were implemented within the AR application. These features allowed on-site personnel to review procedural demonstrations at their own pace, facilitating real-time procedural guidance and post-event reviews for training purposes.
By employing AR to display the results, the system bridged the gap between remote expertise and on-site applications, providing a practical and accessible method for delivering procedural guidance in disaster scenarios.
3. Experimental Evaluation of Remote Support System
Recent studies have shown that perfect spatial alignment in AR-based telementoring systems is not always essential for effective communication. For example, in the ARTEMIS system developed by Gasques et al. 12, alignment errors of 1–2 cm were occasionally observed between virtual annotations and the cadaver viewed by novice trainees. Nevertheless, the expert–novice pair consistently overcame these misalignments through verbal communication, hand gestures, and a shared understanding of anatomical landmarks. These findings suggest that minor positional discrepancies in the AR content can be tolerated in practical clinical scenarios, particularly when accompanied by interactive guidance.
In line with this observation, our system is intended to incorporate real-time voice communication into future versions, which is expected to help resolve minor alignment inconsistencies during remote support. However, to objectively assess the reliability of the current system and quantify its spatial fidelity, we conducted a detailed evaluation of the positional accuracy of AR visualization in this study. Specifically, we evaluated step (6) of the system workflow illustrated in Fig. 1, in which the expert’s avatar was displayed in a real-world environment via AR. To assess the accuracy of this overlay, we measured how far the avatar’s fingertip deviated from the physical cross-shaped markers that had been placed on the debris at the disaster site. This positional offset serves as a quantitative indicator of how precisely the virtual guidance aligns with real-world references.
3.1. Experimental Setup
The purpose of this experiment was to verify the practical feasibility of the proposed remote support system in terms of three key aspects:
-
(1)
Whether scanned disaster environments can be accurately visualized within a VR space.
-
(2)
Whether a motion-captured avatar can interact with the virtual environment.
-
(3)
Whether the results of such interactions can be displayed and spatially evaluated through AR.
This experiment did not aim to reproduce an actual trauma treatment, but rather to test the technical performance of the system in simulating task interactions and presenting positionally accurate results. The experiment was conducted by the author himself.
A mannequin representing a disaster victim was used as the procedural target. Five cross-shaped markers were affixed to critical anatomical locations such as the head, neck, and chest using black vinyl tape (each cross measured approximately 10 mm \(\times\) 50 mm). All markers are shown in Fig. 6. Wooden planks and polystyrene blocks were arranged around the mannequin to simulate collapsed debris. To enhance visual feature detection during scanning and AR localization, a patterned adhesive tape was added to plain polystyrene surfaces in our experimental setup. In practical disaster scenes, however, such auxiliary textures are typically unnecessary; victims’ clothing or hair and debris materials (e.g., wood grains in lumber and cracks in concrete) can provide natural features for both 3D scanning and AR localization.
The entire environment was scanned using a LiDAR-equipped iPhone 16 Pro with the Polycam application. The resulting 3D model was imported into Unity to construct the VR environment. The author, acting as the remote operator, wore a motion capture suit with 50 infrared reflective markers, MANUS METAGLOVES PRO for finger tracking, and a Meta Quest 3 head-mounted display. The task was to align the avatar’s right index fingertip with each of the five target markers.
Each target was touched three times consecutively, and this five-point sequence was repeated in two full sets. The entire session was then repeated, resulting in 60 fingertip contacts.

Fig. 6. Cross markers used in the experiment.
3.2. Positional Accuracy Assessment
To evaluate the spatial precision of AR-based visualization, a photogrammetry-assisted measurement method was employed using RealityCapture j, a commercial software specialized in photogrammetric 3D reconstruction. Initial attempts to measure the error by capturing videos focused on the AR avatar’s fingertip proved unreliable; when only a narrow area around the fingertip was recorded, the number of visible feature points was insufficient for accurate 3D reconstruction, causing misalignment and model fragmentation. In addition, when the virtual fingertip was visually embedded within the mannequin model, RealityCapture failed to reconstruct the penetration depth, making it impossible to accurately measure the positional error. To overcome these limitations, the final approach involved aligning the incomplete 3D model of the AR avatar generated by photogrammetry with a fully accurate avatar model exported from Unity, thereby enabling precise identification of the fingertip position.
Several enhancements were applied to the AR avatar and recording setup to support this process and improve the quality of the photogrammetric reconstruction. During AR playback, the avatar was overlaid onto the real-world scene using a smartphone. To increase the number of identifiable features in the AR projection, a randomized black-and-white texture was applied to the avatar’s surface. Additionally, two circular reference markers were placed beside the debris 20 cm apart, allowing scale calibration within RealityCapture.
The videos were recorded while moving around the AR avatar for one minute to ensure that various angles were captured. The video was then decomposed into 180 frames (one every 20 frames), and the resulting images were input into RealityCapture to generate a 3D point-cloud model of the avatar and surrounding debris. The reconstructed model exhibited sufficient quality for the debris and torso areas of the avatar, although parts such as the arms and fingers were occasionally blurred or duplicated because of motion artifacts.

Fig. 7. Distance measurement in CloudCompare.
Table 1. Measured distance between the avatar’s fingertip and the center of cross markers.
To obtain a clean reference model for the avatar, Unity was used to export the static mesh of the avatar in the exact pose corresponding to each fingertip contact frame. To assist with precise fingertip identification, a 5 mm diameter virtual sphere was placed at the tip of the avatar’s right index finger in Unity before export. The avatar, frozen in the same pose used during AR playback, was then exported from Unity as an .obj file. Both this Unity-generated model and the RealityCapture model were imported into CloudCompare, an open-source software for 3D point clouds and mesh processing. Using this software, eight corresponding landmarks (head, upper back \(\times\)2, right arm \(\times\)2, waist, left arm, and leg) were manually selected for rigid alignment.
Once aligned, the 3D distance between the Unity avatar’s fingertip and the center of each cross-shaped physical marker was measured using CloudCompare. The results are presented in Fig. 7. Five markers were analyzed, each representing a unique anatomical site (e.g., heart, neck, upper limbs, or head), as shown in Table 1. In all five cases, the avatar’s fingertip was visually embedded within the physical marker, indicating a consistent inward displacement of the virtual model relative to the actual contact surface. Fig. 8 shows the distribution of the distance measurements across all the markers. The average positional error across all five points was 27.0 mm, with a standard deviation of 7.0 mm, indicating a moderate variation among the measurements. This level of accuracy suggests that the AR system achieves sufficient alignment precision to visually communicate procedural intent to on-site personnel.

Fig. 8. Boxplot of distance measurements in CloudCompare.
3.3. Limitations
Several limitations of the current system were observed during the experimental sessions.

Fig. 9. Overview of the medical education system.
-
Lack of haptic feedback: Because the system provides no tactile sensation, the avatar’s fingertips sometimes visually penetrated the target surface, particularly during subtle contact attempts.
-
Depth perception challenges: Depth perception occasionally required minor head movements to confirm alignment, but this had minimal impact on task performance.
-
Manual origin alignment: The coordinate origin of the VR scene was determined by the startup pose of Meta Quest 3. This required the operator to physically adjust their position before launching the application to ensure that the virtual victim aligned reasonably well with their intended location.
Although these limitations did not significantly hinder the experimental treatment, they represent areas for improvement in future iterations of the system.
3.4. Technical Contribution
The proposed remote support system offers several technical advantages for field-deployable medical assistance. First, the system architecture enables procedural demonstration and guidance without relying on high-cost or large-scale specialized equipment at the disaster site. By leveraging commonly available consumer devices such as LiDAR-equipped smartphones for 3D scanning and mobile AR-capable tablets for visualization, the system can be rapidly deployed in resource-constrained environments. This design choice enhances practical applicability, as many emergency responders may already possess compatible hardware.
Second, the system incorporates a digital twin approach that allows trauma experts to interact with a three-dimensional reconstruction of a disaster environment through virtual reality. The motion of the expert, captured through optical tracking and finger-sensor gloves, is synchronized with a virtual avatar, enabling an intuitive procedural simulation. This information is then relayed to on-site personnel via augmented reality, to provide spatially aligned visual guidance.
To assess the accuracy of this workflow, a photogrammetry-based evaluation methodology was introduced using RealityCapture. This method enabled numerical measurement of the positional misalignment between the AR-displayed avatar and the physical environment by reconstructing a 3D model from images taken during AR playback. By comparing the position of the avatar’s fingertip and the center of the physical markers in the reconstructed 3D space, the system allowed for objective, three-dimensional error quantification. In this study, the average positional error for the five cross markers and the fingertip was found to be 27.0 mm, demonstrating the feasibility of this approach for evaluating AR spatial fidelity.
4. Medical Education System for Trauma Training
4.1. System Overview
Acquiring trauma surgery skills requires a high level of technical proficiency, which is often difficult to obtain through conventional educational methods such as e-learning or cadaver-based simulations. Moreover, these methods are limited by cost, accessibility, and regulatory constraints.
To address these issues, this study developed a trauma education system that captures expert treatments using a motion capture system and reconstructs them as a digital twin in the Metaverse. This system enables learners to observe and review surgical techniques from arbitrary viewpoints, thereby facilitating self-paced, on-demand training.
An overview of the system is shown in Fig. 9. It combines full-body motion capture, instrument tracking, and a detailed avatar to reproduce procedural movements. In addition, soft-tissue deformation is simulated using position-based dynamics (PBD) 16 to improve the realism of the surgical environment. PBD is a simulation method that updates the positions of particles directly based on constraints rather than calculating forces and accelerations. It is widely used in computer graphics for real-time simulations of soft bodies, clothing, and character physics, because it is stable, fast, and easy to control.
4.2. Expert Demonstration of Trauma Treatments
The experimental setup for the medical education system was partially shared with the remote support system described in Section 2.3, but there were several important differences. The camera used was an optical motion capture system composed of 12 OptiTrack Prime41 cameras, and StretchSense Pro Fidelity Gloves were employed for finger tracking. Unlike the remote support system, this experiment did not involve a HMD because the treatments were performed on a physical trauma simulator visible to the expert.
Unlike the remote support system, the trauma education system was developed using data recorded by a trauma surgeon with specialized experience. An experienced trauma surgeon performed several trauma treatments using surgical instruments and a physical simulator. To enable accurate motion capture, forceps, scalpels, and an electric bone saw were equipped with custom-designed 3D-printed attachments for reflective markers, as shown in Fig. 10. This setup ensured high tracking precision for the instruments.
The recorded trauma treatments included:
-
Amputation
-
Five-point gauze packing
-
Gauze packing for liver injury
-
Pringle maneuver
-
Suture repair of liver injury
-
Temporary abdominal closure
-
Gauze removal
Amputation was assumed to be performed in disaster settings, whereas other treatments were intended to occur at medical relief stations or disaster base hospitals after patient transport. The liver trauma surgery simulator developed by Tojima et al. was used for five-point gauze packing, gauze removal, the Pringle maneuver, and temporary abdominal closure. Training using this simulator was demonstrated to be effective in clinical practice 17.

Fig. 10. Surgical instruments attached with markers.
To enhance the visual and spatial fidelity of the virtual training environment, several physical components were converted into 3D models using Scaniverse k, a LiDAR-based 3D scanning application. These components included a liver trauma surgery simulator, a limb model for amputation training, and an electric bone saw. The resulting 3D models were imported into Unity for integration with the motion capture data and procedural animations.
Reflective metal instruments, such as scalpels and forceps, could not be scanned accurately. To address this issue, a manual modeling method was applied. Each tool was photographed from front and top views. The images were binarized and solidified in Blender, rotated orthogonally, and merged to reconstruct the 3D shape. Boolean operations were then used to generate a final detailed model suitable for simulation.
Currently, reconstructed treatments are visualized and played back within the Unity environment. Viewpoint manipulation is performed manually using mouse controls. Although not yet built as a standalone application, the system is readily portable to iOS, Android, and VR platforms such as Meta Quest, enabling deployment on widely available devices. This flexibility allows the system to be used for clinical training in various environments without requiring specialized hardware.
4.3. Enhancing Realism in Treatment Simulation
To replicate surgical treatments realistically in a virtual environment, the recorded motion data of the expert were applied to an avatar. The avatar was constructed using the same method as described in Section 2.4, utilizing DhaibaWorks to generate a personalized body mesh based on full-body motion capture data and integrating it with skeletal data in Blender.
In addition, to reproduce the opening and closing motion of forceps, the bending angles of the fingers measured by a sensor glove were utilized. Specifically, the angles of the first (DIP), second (PIP), and third (MCP) joints of the middle finger were analyzed, and an algorithm was implemented to transition to the state when the cumulative angle exceeded a predefined threshold. Fig. 11 illustrates the joint states of the fingers during the forceps operation.

Fig. 11. Finger model when using forceps.
The surgical environment was recreated using Unity, incorporating the physics engine Obi Softbody l with PBD for treatments such as gauze packing and gauze removal. PBD is a particle-based physical modeling approach for simulating object deformations. Unlike conventional rigid-body dynamics, PBD applies physical constraints directly, improving computational efficiency. By employing PBD, the system enables the real-time simulation of soft tissue deformations and achieves realistic procedural behavior.
Trauma treatments, including amputation, five-point gauze packing, suture repair of liver injury, and temporary abdominal closure, performed using the system, are shown in Figs. 12(a)–(d), respectively.

Fig. 12. The left images show an expert performing trauma treatments while wearing a motion capture suit. The center and right images respectively show the avatar reproducing the expert’s trauma treatment actions from a third-person and a first-person perspective.
By incorporating this system, learners can utilize the digital twin environment to observe surgical techniques in detail, thereby enhancing their learning experience.

Fig. 13. Error analysis of the digital twin.
4.4. Evaluation of Digital Twin Accuracy
To assess the fidelity of the digital twin environment, we conducted a positional error analysis focusing on the accuracy of tool representation during procedural actions. Specifically, we evaluated the positional accuracy of the virtual forceps tip when contacting predefined markers on the simulator.
During the recording session, circular markers with a diameter of 16 mm were placed in the simulator environment. When the surgeon’s forceps contacted these markers, the corresponding moment in the animation was replayed in Unity, and the 3D distance between the virtual forceps tip and the center of the marker was measured.
The results of this analysis are presented in Fig. 13. The theoretical minimum distance is 8 mm (the radius of the marker). The measured positional error averaged approximately 7 mm, with most samples falling within a \(\pm\)10 mm range. The error was primarily attributed to marker placement precision, motion capture resolution, and 3D model approximation.
This level of error is acceptable for trauma education. The main goal is to understand procedural flow, spatial awareness, and decision-making rather than surgical micromanipulation. The current system demonstrates sufficient spatial accuracy for its intended educational use, with potential for further refinement through improved tracking hardware or higher-fidelity modeling techniques.
5. Discussion
This study presented two complementary systems: a remote support system for trauma treatment in disaster scenarios and a medical education system for trauma treatment training. Both systems combine 3D scanning, motion capture, and immersive visualization to facilitate realistic spatial alignment and the demonstration of medical treatments. This section discusses the system-wide implications, limitations, and future improvements.
5.1. System-Wide Perspective and Practicality
The core operations at disaster or clinical education sites are designed to run on widely available hardware, such as iPhones, iPads, or commercially available VR headsets. This flexibility enables effective response across various disaster sites nationwide.
For instance, disaster-site personnel can create a 3D scan of the environment using a LiDAR-equipped smartphone. This model can then be shared with remote experts, who interact with the reconstructed VR environment. Although the avatar operation via motion capture is conducted in real time, the scanning, data integration, and AR result visualization are not performed in real time and currently require offline processing.
Similarly, in the trauma education system, expert treatments were digitized using motion capture and sensor gloves and integrated into a virtual scene. The physical interaction is visually realistic and computationally efficient, particularly due to the use of PBD. However, the simulation does not yet support complex tissue behaviors such as cutting or rupture.
5.2. Limitations and Future Work
In the remote support system, key limitations include the lack of haptic feedback, which led to some visual inconsistencies (e.g., fingertip penetrating virtual surfaces), and the need for manual alignment due to the dynamic origin setting of the HMD at application startup. Although these limitations did not critically affect operation, they warrant improvement for broader field deployment.
The AR visualization also depends on environmental feature richness. While techniques such as patterned tape improved localization accuracy in the current experiments, additional methods—such as marker-based registration—should be investigated to further stabilize AR overlays.
Although realistic procedural behaviors have been simulated in the medical education system, the use of PBD limits physical accuracy. Future studies should explore finite element method (FEM)-based deformation models to simulate complex surgical interactions and tissue mechanics more precisely 18,19. Additionally, although motion and instrument data from a trauma expert were used, formal evaluations with medical trainees have not yet been conducted. Assessing the system’s educational effectiveness using standardized rubrics, such as OSATS 20,17, remains an important task for future research. Finally, future work will focus on expanding the database of trauma treatments. While this study included a limited number of representative trauma treatments, broader coverage—including various open and minimally invasive techniques—will enable the creation of a comprehensive digital archive for trauma education. This may also be extended to include other domains such as endoscopic or laparoscopic surgery 21, contributing to a wider range of surgical training applications.
These findings highlight the system’s feasibility and applicability for field-based trauma support and clinical education, while also indicating clear paths for technical enhancement and validation.
6. Conclusion
This study developed and evaluated two complementary systems: a remote support system for trauma treatment in disaster scenarios and a trauma education system for medical training. Both systems aim to improve the availability and quality of trauma care and procedural learning through immersive technologies and spatially aligned visualization.
The remote support system enables remote experts to interact with a digital twin of the disaster site, reconstructed from smartphone-based 3D scanning, and to perform simulated treatments in a VR environment. Their movements, captured via motion capture and sensor gloves, are reflected in a digital avatar, whose actions are then conveyed to on-site medical personnel through AR. Experimental results confirmed that the avatar’s interaction with the scanned environment could be visualized in AR with an average positional error of 27.0 mm, demonstrating the system’s feasibility for practical use in field conditions.
The trauma education system was constructed using full-body motion capture data from expert surgeons, capturing realistic procedural behavior with surgical tools. These demonstrations were visualized using a digital avatar in a 3D environment, enabling learners to review complex trauma treatments from multiple perspectives. The positional error between virtual tools and actual reference points was approximately 7 mm, confirming the viability of the system as a precise learning platform.
While promising, several challenges remain. These include the absence of haptic feedback in the VR interface, potential alignment drift in AR visualization, and the limited physical realism of the soft-body simulation based on PBD. Future work will focus on enhancing accuracy through the integration of FEM-based tissue simulation, conducting formal assessments of educational effectiveness using OSATS metrics, and expanding the system’s applicability through the accumulation of a broader trauma procedure dataset, including minimally invasive and endoscopic techniques.
By addressing these limitations and enhancing system capabilities, the proposed approach has the potential to contribute meaningfully to both emergency trauma response and surgical education.
Acknowledgments
This work was supported by Innovative Science and Technology Initiative for Security Grant Number JPJ004596, ATLA, Japan.
- [1] Y. Yanagawa, “Current status of and issues concerning medical relief for huge disasters,” Japanese J. of Neurosurgery, Vol.28, No.9, pp. 561-566, 2019 (in Japanese). https://doi.org/10.7887/jcns.28.561
- [2] C. Yasuda, “Precautions in providing medical treatments beyond physician’s own specialty during emergency medical relief activities and in medical care at evacuation shelters,” The Japanese J. of Quality and Safety in Healthcare, Vol.6, No.2, pp. 252-254, 2011 (in Japanese). https://doi.org/10.11397/jsqsh.6.252
- [3] D. B. Sonshine, A. Caldwell, R. A. Gosselin, C. T. Born, and R. R. Coughlin, “Critically assessing the haiti earthquake response and the barriers to quality orthopaedic care,” Clinical Orthopaedics and Related Research, Vol.470, No.10, pp. 2895-2904, 2012. https://doi.org/10.1007/s11999-012-2333-4
- [4] D. A. Sonshine, A. Caldwell, R. A. Gosselin, C. T. Born, and R. R. Coughlin, “Reply to letter to the editor: Critically assessing the haiti earthquake response and the barriers to quality orthopaedic care,” Clinical Orthopaedics and Related Research, Vol.471, No.2, pp. 692-693, 2013. https://doi.org/10.1007/s11999-012-2685-9
- [5] Z. Materna, M. Španěl, M. Mast, V. Beran, F. Weisshardt, M. Burmester, and P. Smrž, “Teleoperating assistive robots: A novel user interface relying on semi-autonomy and 3D environment mapping,” J. Robot. Mechatron., Vol.29, No.2, pp. 381-394, 2017. https://doi.org/10.20965/jrm.2017.p0381
- [6] J. C. Oliveira, X. Shen, and N. D. Georganas, “Collaborative virtual environment for industrial training and e-commerce,” Proc. of the ACM Symposium on Virtual Reality Software and Technology (VRST), 2000.
- [7] J. J. P. Franko, M. M. Vu, M. E. Parsons, V. Y. Sohn, and J. R. Bingham, “Surgical training for a disaster: Preparation of surgical trainees for victims of conflict,” Military Medicine, Vol.188, Nos.7-8, pp. e2502-e2508, 2023. https://doi.org/10.1093/milmed/usac365
- [8] T. Shichinohe and E. Kobayashi, “Cadaver surgical training in japan: Its past, present, and ideal future perspectives,” Surgery Today, Vol.52, pp. 354-358, 2022. https://doi.org/10.1007/s00595-021-02330-5
- [9] C. M. Harrington, D. O. Kavanagh, J. F. Quinlan, D. Ryan, P. Dicker, D. O’Keeffe, O. Traynor, and S. Tierney, “Development and evaluation of a trauma decision-making simulator in Oculus virtual reality,” The American J. of Surgery, Vol.215, No.1, pp. 42-47, 2018. https://doi.org/10.1016/j.amjsurg.2017.02.011
- [10] G. Huri, M. R. Gülşen, E. B. Karmış, and D. Karagüven, “Cadaver versus simulator based arthroscopic training in shoulder surgery,” Turkish J. of Medical Sciences, Vol.51, No.3, pp. 1179-1190, 2021. https://doi.org/10.3906/sag-2011-71
- [11] K. Ebina, T. Abe, L. Yan, K. Hotta, T. Shichinohe, M. Higuchi, N. Iwahara, Y. Hosaka, S. Harada, H. Kikuchi, H. Miyata, R. Matsumoto, T. Osawa, Y. Kurashima, M. Watanabe, M. Kon, S. Murai, S. Komizunai, T. Tsujita, K. Sase, X. Chen, T. Senoo, N. Shinohara, and A. Konno, “A surgical instrument motion measurement system for skill evaluation in practical laparoscopic surgery training,” PLOS ONE, Vol.19, No.6, Article No.e0305693, 2024. https://doi.org/10.1371/journal.pone.0305693
- [12] D. Gasques, J. G. Johnson, T. Sharkey, Y. Feng, R. Wang, Z. R. Xu, E. Zavala, Y. Zhang, W. Xie, X. Zhang, K. Davis, M. Yip, and N. Weibel, “Artemis: A collaborative mixed-reality system for immersive surgical telementoring,” Proc. of the 2021 CHI Conf. on Human Factors in Computing Systems (CHI’21), Article No.662, 2021. https://doi.org/10.1145/3411764.3445576
- [13] C. Askar and H. Sternberg, “Use of smartphone lidar technology for low-cost 3D building documentation with iPhone 13 pro: A comparative analysis of mobile scanning applications,” Geomatics, Vol.3, No.4, pp. 563-579, 2023. https://doi.org/10.3390/geomatics3040030
- [14] Y. Kako, A. Yoshimura, M. Koyama, N. Miyasato, F. Seki, Y. Nakajima, and F. Satoh, “A study on factors composing rescue-difficulty – An analysis using empirical data on rescue operations at collapsed wooden houses during the 2016 Kumamoto earthquakes –,” J. of Social Safety Science, Vol.36, pp. 65-73, 2020 (in Japanese). https://doi.org/https://doi.org/10.11314/jisss.36.65
- [15] Y. Endo, T. Maruyama, and M. Tada, “Dhaibaworks: A software platform for human-centered cyber-physical systems,” Int. J. Automation Technol., Vol.17, No.3, pp. 292-304, 2023. https://doi.org/10.20965/ijat.2023.p0292
- [16] M. Müller, B. Heidelberger, M. Hennix, and J. Ratcliff, “Position based dynamics,” J. of Visual Communication and Image Representation, Vol.18, No.2, pp. 109-118, 2007. https://doi.org/10.1016/j.jvcir.2007.01.005
- [17] H. Tojima, S. Murakami, S. Poudel, Y. Kurashima, T. Asano, T. Noji, K. Okada, Y. M. Ito, H. Kaneko, Y. Izawa, H. Homma, and S. Hirano, “Development of a simulator and training curriculum for liver trauma surgery training for general surgeons,” Global Surgical Education, Vol.3, Article No.36, 2024. https://doi.org/10.1007/s44186-024-00233-w
- [18] S. Shibuya, N. Shido, R. Shirai, K. Sase, K. Ebina, X. Chen, T. Tsujita, S. Komizunai, T. Senoo, and A. Konno, “Proposal of simulation-based surgical navigation and development of laparoscopic surgical simulator that reflects motion of surgical instruments in real-world,” Int. J. Automation Technol., Vol.17, No.3, pp. 262-276, 2023. https://doi.org/10.20965/ijat.2023.p0262
- [19] X. Chen, D. Sakai, H. Fukuoka, R. Shirai, K. Ebina, S. Shibuya, K. Sase, T. Tsujita, T. Abe, K. Oka, and A. Konno, “Basic experiments toward mixed reality dynamic navigation for laparoscopic surgery,” J. Robot. Mechatron., Vol.34, No.6, pp. 1253-1267, 2022. https://doi.org/10.20965/jrm.2022.p1253
- [20] J. A. Martin, G. Regehr, R. Reznick, H. MacRae, J. Murnaghan, C. Hutchison, and M. Brown, “Objective structured assessment of technical skill (osats) for surgical residents,” The British J. of Surgery, Vol.84, No.2, pp. 273-278, 1997. https://doi.org/10.1046/j.1365-2168.1997.02502.x
- [21] M. Hong, J. W. Rozenblit, and A. J. Hamilton, “Simulation-based surgical training systems in laparoscopic surgery: A current review,” Virtual Reality, Vol.25, pp. 491-510, 2021. https://doi.org/10.1007/s10055-020-00469-z
- [a] “JATEC course,” (in Japanese). https://www.jtcr-jatec.org/index_jatec.html [Accessed May 16, 2025]
- [b] “ASSET course.” https://jp.jssoc.or.jp/modules/info/index.php?content_id=27 [Accessed May 16, 2025]
- [c] “Polycam.” https://poly.cam/ [Accessed May 16, 2025]
- [d] “Unity.” https://unity.com/ [Accessed May 16, 2025]
- [e] “Meta Quest.” https://www.meta.com/jp/en/quest/ [Accessed May 16, 2025]
- [f] “OptiTrack.” https://www.optitrack.com/ [Accessed May 16, 2025]
- [g] “Metagloves Pro.” https://www.manus-meta.com/products/metagloves-pro [Accessed May 16, 2025]
- [h] “Blender.” https://www.blender.org/ [Accessed May 18, 2025]
- [i] “Immersal.” https://immersal.com/ [Accessed May 18, 2025]
- [j] “RealityCapture,” (in Japanese). https://www.capturingreality.com/ [Accessed May 18, 2025]
- [k] “Scaniverse,” (in Japanese). https://scaniverse.com/ [Accessed May 18, 2025]
- [l] “Obi Softbody.” https://assetstore.unity.com/packages/tools/physics/obi-softbody-130029 [Accessed May 18, 2025]
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.