single-jc.php

JACIII Vol.30 No.4 pp. 1258-1264
(2026)

Research Paper:

Gaze Analysis System for Emotional Attunement in Human–Robot Interactions

Yuka Sone* and Jinseok Woo**,† ORCID Icon

*Sustainable Engineering Program, Graduate School of Engineering, Tokyo University of Technology
1404-1 Katakuramachi, Hachioji, Tokyo 192-0982, Japan

**Department of Mechanical Engineering, School of Engineering, Tokyo University of Technology
1404-1 Katakuramachi, Hachioji, Tokyo 192-0982, Japan

Corresponding author

Received:
December 19, 2025
Accepted:
March 14, 2026
Published:
July 20, 2026
Keywords:
gaze analysis, mixed reality, emotional attunement, human–system interaction
Abstract

Recent advances in artificial intelligence and robotics have accelerated the deployment of service robots in daily environments. However, many systems still lack adaptive responsiveness to the nonverbal behaviors of users. This study proposes a gaze-based user analysis system integrated into a mixed reality (MR) smart home environment to support attentional intention-aware human–system interactions. Rather than directly estimating emotional states, the proposed approach infers attentional intentions of users based on gaze behavior as an operational proxy for emotional attunement. Using gaze data collected through HoloLens 2, we develop a machine learning model based on a long short-term memory network combined with a mixture density network to predict future gaze coordinates in a three-dimensional space. The predicted gaze information is shared with a robotic partner to enable proactive context-aware information support. The proposed system demonstrates the feasibility of leveraging gaze prediction to anticipate user focus and provide adaptive support in MR-based smart environments.

Gaze analysis in an MR smart home system

Gaze analysis in an MR smart home system

Cite this article as:
Y. Sone and J. Woo, “Gaze Analysis System for Emotional Attunement in Human–Robot Interactions,” J. Adv. Comput. Intell. Intell. Inform., Vol.30 No.4, pp. 1258-1264, 2026.
Data files:

1. Introduction

Recently, owing to the rapid advancement of artificial intelligence (AI) and robotics technologies, the range of services that robots can perform in daily human life has significantly expanded 1. For example, robots that provide services traditionally performed by humans, such as caregiving robots in home environments, guide robots at tourist sites, and customer service robots in stores and restaurants, are rapidly becoming integrated into everyday life 2,3,4. These service robots are drawing considerable attention as promising solutions to labor shortages in aging societies 5,6,7.

However, despite the widespread adoption of these service robots, users often report difficulties in perceiving human warmth or empathy in the services provided by robots 8,9. This issue is considered not merely a matter of functionality, but rather a consequence of the fact that current service robots tend to engage in mechanical and one-directional interactions, often failing to fully comprehend the emotional context of human users 10.

Tomasello explained the origins of human language in terms of the development of cooperative sociality and shared intentionality, arguing that the essence of language lies in the capacity for mutual understanding and coordination toward shared goals 11,12. In particular, emotional attunement enables interlocutors to infer the affective state of speakers and thereby achieve a more accurate understanding of their internal condition. Accordingly, emotional attunement can be regarded as a key component of not only human–human interactions but also realizing natural communication in human–robot and human–system interactions. In this study, emotional attunement is not treated as direct emotion recognition. Instead, we define emotional attunement through the inference of user attentional intention derived from gaze behavior. Because gaze reflects attentional focus and cognitive engagement, predicting gaze trajectories enables the system to estimate the attentional focus of a user in real time. This attentional intention is used as a practical proxy for emotionally attuned interactions, allowing the system to adapt its responses based on the anticipated user focus rather than explicit emotional labels. By adopting this formulation, the present study emphasizes an engineering-oriented approach to nonverbal interactions, focusing on attention-aware support rather than affective state estimation.

In human–human interactions, emotional attunement is facilitated by the continuous interpretation of multiple nonverbal cues 13. Individuals indirectly infer the internal state of another person by observing subtle changes in facial expressions, shifts in gaze direction, variations in posture, and dynamic gestures that accompany speech or contextual behavior. These nonverbal signals provide rich information on affective intent, engagement levels, and underlying psychological conditions, enabling interlocutors to construct a more precise understanding of the emotional state of the other person. This multimodal information plays a crucial role in maintaining coordinated interactions and sustaining mutual understanding. Building on the view that emotional attunement in human communication is achieved through the continuous interpretation of nonverbal cues, this study positions gaze behavior as a particularly informative channel for extending these mechanisms to service-oriented robotic systems, as illustrated in Fig. 1.

Because gaze reflects the attentional focus, cognitive processing, and affective engagement of a user, analyzing gaze patterns offers a promising means to infer the internal state of the user in contexts where direct verbal expressions may be limited or ambiguous. Therefore, we developed a range of robotic systems that leverage nonverbal information, particularly gaze-based cues, to facilitate more adaptive, context-aware, and emotionally attuned interactions with users 14,15,16. Building on these prior developments, the present study introduces a gaze-based analytic approach designed to enhance the ability of robot systems to interpret the internal states of users for emotionally attuned interactions.

The remainder of this paper is organized as follows. Section 2 describes the proposed systems used in this study. Section 3 describes the structure of the machine learning system. Section 4 presents an analysis of the experimental results. Section 5 discusses directions for future work.

figure

Fig. 1. Conceptual framework of gaze-based user analysis.

2. Overall System Configuration

In this study, we employ a smart home system as an interactive platform. This system is designed for domestic environments and supports remote operations. We use HoloLens 2 17, a head-mounted display capable of mixed reality (MR) interactions, as the input device for interacting with the smart home system. By incorporating the MR technology, the system can utilize digital twin representations, allowing users in remote locations to visually access real-time information on the operation of their smart home devices.

To ensure applicability across diverse residential environments, the proposed system controls home appliances using attachable modules that can be retrofitted into general household settings. The target appliances include curtains and lighting that are frequently used in daily life, and the system additionally provides access to the internal state of a refrigerator 7. By monitoring the operational states of these appliances, we can make a simple estimation of the lifestyle patterns of residents within the environment.

figure

Fig. 2. Overview of the system flow.

figure

Fig. 3. Physical robot partner system.

The overall system flow is illustrated in Fig. 2. When input device (1) HoloLens 2 is activated, (2) information about the actual home appliances is retrieved from the database, and the virtual appliances displayed in the HoloLens 2 environment are updated to match their physical counterparts. When the user issues an operation command through HoloLens 2, the command is transmitted to the database, and (3) a robot or ESP32-S3 accesses the database to operate the corresponding physical appliances. Both HoloLens 2 and ESP32-S3 periodically communicate with the database, enabling real-time remote interactions 18.

In this system, the robot functions as an information mediator for users in the physical environment, as illustrated in Fig. 3, by notifying them of appliance activation events and providing additional appliance-related information 19. The virtual environment displayed in the HoloLens 2 is constructed based on the physical environment in which the appliances are installed, as shown in Fig. 4. The basic behaviors of the virtual robot partner are illustrated in Fig. 5. To enable remote control within HoloLens 2, the system measures both the hand movements and gaze behavior of the users. The users can operate all the appliances using a combination of hand gestures and gaze input. In particular, gaze information is used not only for appliance manipulation but also for enabling the robot to provide informational support.

Gaze information is obtained using two infrared sensors integrated into HoloLens 2. As shown in Fig. 6, the gaze data are collected and analyzed, and the point of gaze contact is indicated by a red dot. After completing the calibration procedure, the system accurately measures the gaze of each user. The gaze point is computed by identifying the position at which the two independently captured gaze rays from the left and right eyes are closest to each other. To determine the gaze point of the user within the virtual environment, the gaze coordinates obtained from the local head coordinate system of the head-mounted display are transformed into the world coordinate system.

figure

Fig. 4. Examples of the physical and virtual environments used in experiments.

figure

Fig. 5. Basic behaviors of the virtual robot partner.

figure

Fig. 6. Illustration of gaze measurement with HoloLens 2.

3. Machine Learning Model

In this study, a machine learning model is designed to predict future gaze points to infer user attentional intention rather than directly estimating emotional states. By analyzing time-series gaze data, the system estimates the forthcoming focus of attention of a user that is used as an operational proxy for emotional attunement in human–system interactions. This prediction enables the robot partner to anticipate the user interests and provide adaptive, context-aware information support.

To achieve this, we employ a machine learning framework that combines a long short-term memory network with a mixture density network (LSTM–MDN) to model the temporal dynamics and uncertainty of gaze behavior.

figure

Fig. 7. Architecture of the machine learning model.

The overall architecture of the model is illustrated in Fig. 7. The model aims to predict future gaze point \(\mathbf{g}_{K+1} \in \mathbb{R}^3\) based on a sequence of past observations. This study uses a dataset consisting of 9,233 gaze sequences in which each sequence represents a set of gaze data points sampled at 0.3 s intervals. The dataset is divided into training, validation, and testing sets at ratios of 70%, 15%, and 15%, respectively. Let \(\mathbf{g}_t = (x_t, y_t, z_t)^\top\) denote the gaze coordinates at timestep \(t \in \{1, \dots, K\}\), where \(\top\) denotes the transpose operator and \(K=10\) is the sequence length. To capture the dynamics of gaze behavior, we compute velocity \(\mathbf{v}_t = (v_{x,t}, v_{y,t}, v_{z,t})^\top\) and acceleration \(\mathbf{a}_t = (a_{x,t}, a_{y,t}, a_{z,t})^\top\) using finite differences with a sampling interval of \(s = 0.3\) s as follows:

\begin{equation} \label{eq:preprocessing} \mathbf{v}_t = \frac{\mathbf{g}_t - \mathbf{g}_{t-1}}{s}, \quad \mathbf{a}_t = \frac{\mathbf{v}_t - \mathbf{v}_{t-1}}{s}. \end{equation}
The input vector at each timestep is constructed as a concatenated feature vector \(\mathbf{u}_t = (x_t, y_t, z_t, v_{x,t}, v_{y,t}, v_{z,t}, a_{x,t}, a_{y,t}, a_{z,t})^\top \in \mathbb{R}^9\). The model receives a sequence of consecutive timesteps \(\mathbf{U} = (\mathbf{u}_1, \dots, \mathbf{u}_K)\) to perform the prediction.

To extract the temporal dependencies from input sequence \(\mathbf{U}\), we employ an LSTM network. Recurrent computations within the LSTM unit at each timestep \(t\) are governed by the following equations.

\begin{equation} \label{eq:lstm} \begin{cases} \mathbf{i}_t = \sigma\left(\mathbf{W}_{iu}^{\left(L\right)} \mathbf{u}_t + \mathbf{b}_{iu}^{\left(L\right)} + \mathbf{W}_{hu}^{\left(L\right)} \mathbf{h}_{t-1} + \mathbf{b}_{hu}^{\left(L\right)}\right), \\ \mathbf{f}_t = \sigma\left(\mathbf{W}_{fu}^{\left(L\right)} \mathbf{u}_t + \mathbf{b}_{fu}^{\left(L\right)} + \mathbf{W}_{hf}^{\left(L\right)} \mathbf{h}_{t-1} + \mathbf{b}_{hf}^{\left(L\right)}\right), \\ \tilde{\mathbf{c}}_t = \tanh\left(\mathbf{W}_{gu}^{\left(L\right)} \mathbf{u}_t + \mathbf{b}_{gu}^{\left(L\right)} + \mathbf{W}_{hg}^{\left(L\right)} \mathbf{h}_{t-1} + \mathbf{b}_{hg}^{\left(L\right)}\right), \\ \mathbf{o}_t = \sigma\left(\mathbf{W}_{ou}^{\left(L\right)} \mathbf{u}_t + \mathbf{b}_{ou}^{\left(L\right)} + \mathbf{W}_{ho}^{\left(L\right)} \mathbf{h}_{t-1} + \mathbf{b}_{ho}^{\left(L\right)}\right), \\ \mathbf{c}_t = \mathbf{f}_t \odot \mathbf{c}_{t-1} + \mathbf{i}_t \odot \tilde{\mathbf{c}}_t, \\ \mathbf{h}_t = \mathbf{o}_t \odot \tanh\left(\mathbf{c}_t\right), \end{cases} \end{equation}
where \(\mathbf{i}_t, \mathbf{f}_t, \mathbf{o}_t\), and \(\tilde{\mathbf{c}}_t\) represent the input gate, forget gate, output gate, and cell candidate, respectively. Weight matrices \(\mathbf{W}^{(L)}\) and bias vectors \(\mathbf{b}^{(L)}\) are the trainable parameters of the LSTM layer. Final hidden state \(\mathbf{h}_K\) is subsequently fed into the MDN layer.

To account for the multi-modalnature and inherent uncertainty of gaze movements, the MDN maps hidden state \(\mathbf{h}_K\) to the parameters of a Gaussian mixture model (GMM) with \(M=5\) components. The MDN estimates mixing coefficients \(\pi_m\), means \(\boldsymbol{\mu}_m = (\mu_{m,x}, \mu_{m,y}, \mu_{m,z})^\top\), and standard deviations \(\boldsymbol{\sigma}_m = (\sigma_{m,x}, \sigma_{m,y}, \sigma_{m,z})^\top\) for each component \(m \in \{1, \dots, M\}\) using MDN-specific parameters \(\mathbf{W}^{(M)}\) and \(\mathbf{b}^{(M)}\):

\begin{equation} \label{eq:mdn_params} \begin{cases} \boldsymbol{\pi} = \text{softmax}\left(\mathbf{W}_{\pi}^{\left(M\right)} \mathbf{h}_K + \mathbf{b}_{\pi}^{\left(M\right)}\right), \\ \boldsymbol{\mu}_m = \mathbf{W}_{\mu,m}^{\left(M\right)} \mathbf{h}_K + \mathbf{b}_{\mu,m}^{\left(M\right)}, \\ \boldsymbol{\sigma}_m = \exp\left(\mathbf{W}_{\sigma,m}^{\left(M\right)} \mathbf{h}_K + \mathbf{b}_{\sigma,m}^{\left(M\right)}\right). \end{cases} \end{equation}
The conditional probability density of next gaze point \(\mathbf{g}_{K+1}\) is defined as a weighted sum of \(M\) Gaussian distributions.
\begin{equation} \label{eq:gmm} P\left(\mathbf{g}_{K+1} \mid \mathbf{h}_K\right) = \sum_{m=1}^{M} \pi_m \mathcal{N}\left(\mathbf{g}_{K+1} \mid \boldsymbol{\mu}_m, \boldsymbol{\sigma}_m\right). \end{equation}
Assuming a diagonal covariance matrix, each component \(\mathcal{N}\) is calculated by breaking down the product as follows:
\begin{align} \label{eq:gaussian_component} \mathcal{N}\left(\mathbf{g}_{K+1} \mid \boldsymbol{\mu}_m, \boldsymbol{\sigma}_m\right) &= \prod_{d \in \{x,y,z\}} \frac{1}{\sqrt{2\pi}\sigma_{m,d}} \notag\\ &\phantom{=~}\times \exp\left( -\frac{\left(d_{K+1} - \mu_{m,d}\right)^2}{2\sigma_{m,d}^2} \right). \end{align}
The model was implemented using PyTorch and trained by minimizing the mean negative log-likelihood loss function over a batch of size \(N_{\mathit{batch}}\).
\begin{equation} \label{eq:loss} \mathcal{L} = - \frac{1}{N_{\mathit{batch}}} \sum_{n=1}^{N_{\mathit{batch}}} \log \left( P\left(\mathbf{g}_{K+1}^{(n)} \mid \mathbf{h}_K^{(n)}\right) \right). \end{equation}
We used the Adam optimizer with a learning rate of \(1 \times 10^{-3}\) and a batch size of 32. To prevent overfitting and ensure convergence, we employed early stopping with a patience of 15 epochs, and a learning rate scheduler (ReduceLROnPlateau). The hidden size was determined through a grid search within the range of \(\{32, 64, 128\}\) based on the validation performance.

4. Experimental Results

In the present study, the experimental evaluation intentionally focused exclusively on gaze information to isolate attentional intention as the primary nonverbal cue. Although emotional communication in human interactions is inherently multimodal, incorporating facial expressions, speech, and body gestures, the purpose of this experiment was to examine the feasibility of attention-aware support based solely on gaze behavior. Multimodal affective analyses and comparative evaluations with and without gaze information are beyond the scope of the current study and are reserved for future studies.

4.1. Experimental Overview

Based on this design choice, we conducted an experiment to perform real-time gaze prediction using the constructed learning model. Real-time gaze coordinates were transmitted from HoloLens 2 to the server that processed the input data to predict the results. In this system, gaze measurement data were sent to the server every 0.3 s, and the system yielded predictions of gaze coordinates 0.3 s into the future. The 0.3-s interval was selected by considering the average duration of human eye blinks and was thus adopted as the gaze measurement period in the system. Note that the measurement samples in which the gaze coordinates were located extremely far from the virtual environment displayed in the virtual space were excluded from use.

To obtain the results as accurately as possible, the experiment was conducted in a quiet environment without any objects obstructing the field of view of the participant. During the experiment, each of the three appliances targeted for operation—refrigerator, curtain, and lighting—was used at least once. Because the system performed prediction every 0.3 s, the duration of each experiment was set to at least 30 s.

As shown in Fig. 8, the experimental procedure consisted of (1) a tutorial explaining the method of using the system followed by (2) the experiment itself. Finally, (3) the experimental results were computed and the prediction accuracy was evaluated. Fig. 9 shows the experimental setup in action. Specifically, (a) presents an external view of the user during the experiment, and (b) displays the system interface from the user perspective.

figure

Fig. 8. Overview of the experimental procedure.

figure

Fig. 9. Experimental environment.

4.2. Gaze Prediction Performance

The experiment was conducted with 14 participants in their twenties. For each experiment, the deviation between the measured and predicted gaze coordinates was calculated. The averaged results are shown in Fig. 10; the horizontal axis represents the error distance between the predicted and measured coordinates, and the vertical axis represents the cumulative ratio.

To assess the overall predictive performance, we reported the 70th percentile error that represented the error distance within which 70% of all predictions fell. This percentile-based metric was adopted to robustly characterize the typical prediction accuracy under the multimodal distribution modeled by the MDN. A percentile-based evaluation is particularly suitable for gaze prediction in which error distributions are often skewed and multimodal.

As shown in Fig. 10, the 70th percentile error in a three-dimensional space was 0.29 m. To illustrate the scale of this distance, the size of the virtual environment constructed in the system is shown in Figs. 11 and 12. Fig. 11 shows the virtual environment from the front in which the room width and height are 0.35 m and 0.24 m, respectively. Fig. 12 shows the environment from above in which the room depth is 0.56 m. These results indicated that an error of 0.29 m in a three-dimensional space corresponded to 18.5% of the maximum possible error.

figure

Fig. 10. Experiment result of the 3D space.

figure

Fig. 11. Front view of the virtual environment.

figure

Fig. 12. Top view of the virtual environment.

Next, we considered the error in two-dimensional planes. The average errors in each plane computed from the experimental results are shown in Fig. 10. As shown in Fig. 13, on the \(X\)\(Y\) plane, the 70th percentile error was 0.09 m, and the corresponding values on the \(Y\)\(Z\) and \(X\)\(Z\) planes were 0.26 m and 0.29 m, respectively. These results indicated that the errors in the depth direction (\(Z\)-axis) were relatively large, whereas the errors in the horizontal \(X\)- and vertical \(Y\)-directions were smaller.

figure

Fig. 13. Experiment result of the 2D space: \(X\)\(Y\), \(Y\)\(Z\), and \(X\)\(Z\).

One factor contributing to the larger error in the depth direction was the perceptual error in recognizing the distances within the virtual space 20,21. As revealed in the observational evaluation, participants often perceived the buttons to be located closer than their actual positions when attempting to press them. This suggested that the human perceptual error in the \(Z\)-direction was relatively large, which resulted in the observed discrepancy between the predicted and measured values.

Despite the larger prediction error observed along the depth, that is \(Z\)-axis, the proposed system remained capable of supporting intention-level interactions because appliance discrimination in a smart home environment primarily relied on spatial separation in the \(X\)\(Y\) plane. In practice, the predicted gaze information is used to estimate the general user focus rather than precise depth positioning, allowing the system to tolerate \(Z\)-axis uncertainty while still providing effective context-aware assistance.

This study predicted gaze coordinates 0.3 s into the future. However, for more adaptive attention-aware interaction, the prediction horizon must be extended beyond 0.3 s to longer time intervals. Therefore, extending the prediction horizon (for example, to 0.9 s and 1.5 s) is considered important for future personalization and will be addressed in future studies.

4.3. Application Example of Gaze-Based Adaptive Information Support

To illustrate using predicted gaze information in practical interactions, this subsection presents an application example of gaze-based adaptive information support using a robotic partner. Within the system, a robot provided information and proposed services to a user. This robot introduced home appliances located in the direction of the gaze of the user. By tuning this function according to the results of gaze prediction, the system could offer services that were better tailored to each user.

Table 1. Information provided by the robot related to home appliances.

figure
figure

Fig. 14. Adaptive information provision by the robot using gaze prediction and schedule.

As a basic form of information provision, as presented in Table 1, each home appliance provided various types of information based on the current time. In addition to these gaze prediction results, we believed that more personalized information could be provided by referring to the routine-based behavioral schedule of a user in a real environment. For example, as illustrated in Fig. 14, when the predicted gaze position was near the refrigerator and the current time was 7:00 AM, the schedule indicated breakfast time. In this case, the robot was expected to provide information regarding the ingredients in the refrigerator that were typically used for breakfast, thereby realizing a service adapted to the situation of the user.

5. Conclusion and Future Works

In this study, we proposed a gaze-based user analysis system integrated into an MR smart home environment to support attentional intention-aware human–system interactions. Rather than directly estimating emotional states, the system inferred the attentional intentions of users from their gaze behavior as an operational proxy for emotional attunement, thereby enabling adaptive and context-aware information support.

The proposed system enabled remote operation of home appliances through a virtual environment while continuously measuring gaze information to predict future gaze coordinates 0.3 s ahead using an LSTM–MDN model. Experiments conducted with 14 participants demonstrated that 70% of predictions fell within an error of 0.29 m in three-dimensional space, with notably higher accuracy on the \(X\)\(Y\) plane. These results indicated that gaze-based attentional intention prediction could be achieved with sufficient accuracy for interactive assistance; however, depth-related uncertainty remains a limitation that requires further consideration. By sharing the predicted gaze information with a robotic partner, the system could anticipate user focus and provide proactive situation-aware support. This approach demonstrated the feasibility of leveraging gaze prediction for attention-aware interactions in MR-based smart environments.

Future work will extend the prediction horizon beyond 0.3 s, incorporate additional nonverbal modalities, and evaluate the system with broader user populations, including elderly participants, to improve generalizability. Furthermore, adaptive service strategies that integrate predicted gaze with contextual information such as time and user routines will be explored to enhance proactive assistance, interaction efficiency, and context-aware system responses in real-world smart home scenarios. Furthermore, adaptive service strategies that integrate predicted gaze with contextual information such as time and user routines will be explored to enhance proactive assistance, interaction efficiency, and context-aware system responses in real-world smart home environments.

Acknowledgments

This work was supported by JSPS KAKENHI (Grant-in-Aid for Early-Career Scientists) (Grant Number 23K17261).

References
  1. [1] M. C.-T. Tai, “The impact of artificial intelligence on human society and bioethics,” Tzu Chi Medical J., Vol.32, No.4, pp. 339-343, 2020. https://doi.org/10.4103/tcmj.tcmj_71_20
  2. [2] T. Fong, I. Nourbakhsh, and K. Dautenhahn, “A survey of socially interactive robots,” Robotics and Autonomous Systems, Vol.42, Nos.3-4, pp. 143-166, 2003. https://doi.org/10.1016/S0921-8890(02)00372-X
  3. [3] C. Bartneck, T. Kanda, O. Mubin, and A. Al Mahmud, “Does the design of a robot influence its animacy and perceived intelligence?,” Int. J. of Social Robotics, Vol.1, No.2, pp. 195-204, 2009. https://doi.org/10.1007/s12369-009-0013-7
  4. [4] G. Bardaro, A. Antonini, and E. Motta, “Robots for elderly care in the home: A landscape analysis and co-design toolkit,” Int. J. of Social Robotics, Vol.14, No.3, pp. 657-681, 2022. https://doi.org/10.1007/s12369-021-00816-3
  5. [5] L. C. Cesário, P. Barbosa, P. A. C. Miguel, and G. H. Mendes, “Service robots in caring for older adults: Uncovering the current conceptual and intellectual structures and future research agenda,” Archives of Gerontology and Geriatrics, Vol.131, Article No.105755, 2025. https://doi.org/10.1016/j.archger.2025.105755
  6. [6] J. Abdi, A. Al-Hindawi, T. Ng, and M. P. Vizcaychipi, “Scoping review on the use of socially assistive robot technology in elderly care,” BMJ Open, Vol.8, No.2, Article No.e018815, 2018. https://doi.org/10.1136/bmjopen-2017-018815
  7. [7] Y. Sone, C. Matsumoto, J. Woo, and Y. Ohyama, “Development of a control support system for smart homes using the analysis of user interests based on mixed reality,” ROBOMECH J., Vol.12, Article No.2, 2025. https://doi.org/10.1186/s40648-024-00288-w
  8. [8] I. Leite, A. Pereira, S. Mascarenhas, C. Martinho, R. Prada, and A. Paiva, “The influence of empathy in human–robot relations,” Int. J. of Human-Computer Studies, Vol.71, No.3, pp. 250-260, 2013. https://doi.org/10.1016/j.ijhcs.2012.09.005
  9. [9] M. Coeckelbergh, “Artificial companions: Empathy and vulnerability mirroring in human-robot relations,” Studies in Ethics, Law, and Technology, Vol.4, No.3, 2010. https://doi.org/10.2202/1941-6008.1126
  10. [10] M. Spezialetti, G. Placidi, and S. Rossi, “Emotion recognition for human-robot interaction: Recent advances and future perspectives,” Frontiers in Robotics and AI, Vol.7, Article No.532279, 2020. https://doi.org/10.3389/frobt.2020.532279
  11. [11] M. Tomasello, “Origins of Human Communication,” MIT Press, 2010.
  12. [12] M. Tomasello, “A Natural History of Human Thinking,” Harvard University Press, 2014. https://doi.org/10.4159/9780674726369
  13. [13] S. Balzarotti, L. Piccini, G. Andreoni, and R. Ciceri, ““I know that you know how I feel”: Behavioral and physiological signals demonstrate emotional attunement while interacting with a computer simulating emotional intelligence,” J. of Nonverbal Behavior, Vol.38, No.3, pp. 283-299, 2014. https://doi.org/10.1007/s10919-014-0180-6
  14. [14] C. Matsumoto, Y. Sone, J. Woo, and Y. Ohyama, “Development of a gesture analysis system for human system interaction,” 2024 Joint 13th Int. Conf. on Soft Computing and Intelligent Systems and 25th Int. Symp. on Advanced Intelligent Systems (SCIS&ISIS), 2024. https://doi.org/10.1109/SCISISIS61014.2024.10760070
  15. [15] Y. Sone and J. Woo, “Design of a human-centric robotic system for user support based on gaze information,” J. Adv. Comput. Intell. Intell. Inform., Vol.29, No.4, pp. 796-802, 2025. https://doi.org/10.20965/jaciii.2025.p0796
  16. [16] J. Woo and J. Hu, “System for analyzing user interest based on eye gaze responses to enhance empathy with users,” J. Adv. Comput. Intell. Intell. Inform., Vol.29, No.3, pp. 641-648, 2025. https://doi.org/10.20965/jaciii.2025.p0641
  17. [17] P. Balakrishnan and H.-J. Guo, “Hololens 2 technical evaluation as mixed reality guide,” Int. Conf. on Human-Computer Interaction, pp. 145-165, 2024. https://doi.org/10.1007/978-3-031-61041-7_10
  18. [18] C. Matsumoto, Y. Sone, J. Woo, and Y. Ohyama, “Development of a user-friendly interface for a smart home system based on IoT technology,” The 20th World Congress of the Int. Fuzzy Systems Association (IFSA 2023), 2023.
  19. [19] J. Woo, Y. Ohyama, and N. Kubota, “Robot partner development platform for human-robot interaction based on a user-centered design approach,” Applied Sciences, Vol.10, No.22, Article No.7992, 2020. https://doi.org/10.3390/app10227992
  20. [20] J. W. Kelly, T. A. Doty, M. Ambourn, and L. A. Cherep, “Distance perception in the Oculus Quest and Oculus Quest 2,” Frontiers in Virtual Reality, Vol.3, Article No.850471, 2022. https://doi.org/10.3389/frvir.2022.850471
  21. [21] C. S. Rosales, G. Pointon, H. Adams, J. Stefanucci, S. Creem-Regehr, W. B. Thompson, and B. Bodenheimer, “Distance judgments to on-and off-ground objects in augmented reality,” 2019 IEEE Conf. on Virtual Reality and 3D User Interfaces (VR), pp. 237-243, 2019. https://doi.org/10.1109/VR.2019.8798095

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Jul. 19, 2026