The combination of BOB and PERCEPTION NEURON helps to conduct large-scale assessments of psychological stress and biomechanical strain in manufacturing environments.
Release time:
2022-12-12 14:14
Source:
Using PERCEPTION NEURON and BOB as tools
Large-scale assessment of psychological stress and biomechanical strain in manufacturing environments
Lucas Paletta, Harald Ganster, Michael Schneeberger, Martin Pszeida, Gerald Lodron;
JOANNEUM RESEARCH, Graz, Austria
Katrin Pechstädt, Michaela Spitzer, Christiane Reischl;
FH JOANNEUM, Bad Gleichenberg / Graz, Austria
Abstract:
In the future manufacturing industry, human-machine interaction will develop towards flexible and intelligent collaboration. It will meet the demands for optimizing assembly processes and for active and skilled human behavior. Recently, substantial progress has been made in human factors engineering through detailed task analysis. However, there is still a lack of precise measurements of cognitive and sensorimotor patterns to analyze long-term mental and physical stress. This experiment proposes a new method that can measure cognitive load in real-time based on functional analysis and biomechanical strain from non-invasive wearable sensors. The method uses precise stereometric measurement equipment to recover 3D information of the working subject. The worker is equipped with eye-tracking glasses and a set of wearable accelerometers. Wireless connections transmit sensor data to a nearby PC for monitoring. Data analysis then recovers the 3D geometric structure of gaze and observes the frustum of the working subject while further extracting its task-switching rate and assessing its biomechanical strain approximations based on skeletons in different postures. AI-based assessments indicate that the first results closely match the activity analysis results conducted by occupational therapists.
Introduction:
The production industry is currently undergoing a continuous transition towards the Fourth Industrial Revolution, creating virtual replicas of the physical world and decentralized decision-making through modular structural entities and cyber-physical systems monitoring physical processes, thus nurturing "smart factories." Recently, substantial progress has been made in human factors engineering through detailed task analysis. However, the extraction of task descriptions still relies on manual analysis and detailed, time-consuming descriptions of human-machine interaction monitoring based on video. In this way, the long-term observation and analysis of work is too costly, which means there is still a lack of precise measurements of cognitive and sensorimotor patterns for analyzing long-term mental and physical stress.
From the perspective of service quality, attention to the gains and losses in productivity plays an important role. Attention metrics shift to measurements of concentration, cognitive load, and situational awareness, as well as early indicators of fatigue. The quality of attention is directly related to the quality of decision-making, thus relevant service quality (e.g., increased fatigue) indicates that the safety of decision-making is reduced due to shortened attention duration, reduced reaction time, and decreased accuracy. It is well known that undetected and therefore frequently occurring fatigue leads to more team misunderstandings, reduced morale and motivation, as well as increased short-term and long-term absenteeism and medical investments due to related physical illnesses. Therefore, the quality of attention is related to avoiding productivity loss due to distraction, errors, inability to concentrate, and lack of motivation.
Our work aims to achieve challenging long-term goals in order to enable unobtrusive long-term analysis of mental, physical, emotional, and motivational stress of manufacturing site workers. We believe that purely observation-based analysis cannot fully explain the psychophysical processes that reflect long-term mental and physical stress. The technical challenge of this work is to measure the local infrastructure that determines the geometric and functional conditions of worker interaction, as well as to describe the psychophysiological and biomechanical parameters of human states, intentions, and activities.
This experiment proposes a new method that can measure cognitive load and biomechanical strain in real-time using non-invasive wearable sensors. In particular, cognitive load is calculated based on eye movement analysis, while eye movements are derived from the semantics of the recovered 3D structure. From these observations, we generate several metrics related to "task switching," which are based on the typical characteristics of attention-driven task behavior. Finally, our experimental measures can provide better approximations of human states, performance, and strain parameters than external observations.

Figure 1: Wearable data collection for large-scale assessment of mental load and biomechanical strain in factory workplaces. The selection and placement of workplaces demonstrate the challenging requirements for cognitive and biomechanical strain.
Related Work:
In the field of human-machine interaction, human-centered variables have played an important role for quite some time. Steinfeld et al. [1] describe parameters such as situational awareness, workload, and mental models using human-machine interaction as an example. These human-related variables are crucial for assessing the interaction metrics between people. Human factors are essential as industrial robots enable human and robotic workers to work side by side as collaborators and assess user experiences with robots while understanding how humans feel during the interaction process with robots (Figure 1). As Huang and Mutlu [2] describe, attention and gaze provide a means of action prediction and intention, allowing for seamless and efficient collaboration with human peers. Similarly, machines must rely on predictions of human workers' behaviors, emotions, task-specific actions, and intentions to plan their actions, for example, using human-in-the-loop architectures [2] for anticipatory control, enabling robots to actively perform task actions based on observed gaze patterns, thus predicting the actions of their human partners.
Measuring and simulating human attention, focus, and situational awareness based on eye movements and corresponding triggering information recovery is essential for understanding human action planning, mental workload, and the service quality achieved as a result. Gaze as an input device has proven beneficial, for example, in workshop management [3]. Finally, Carnegie Mellon University and Admoni et al. [4] introduced gaze-based cyber-physical control as an additional asset. Gaze seems to have broad opportunities to provide information for monitoring, assessment, interactive devices, and action prediction.
The work most relevant to the method proposed in this experiment is Santner et al. [5], who performed 3D gaze recovery using eye-tracking glasses and SLAM methods for 3D information recovery. The proposed method builds on the results of that work but essentially extends it, applying more advanced handheld devices for better infrastructure recovery and adding a complete method for extracting information about executive functions, as well as estimates related to biomechanical strain, for comprehensive assessment of strain in industrial work cells through human-machine interaction.
Principles and Data Processing of Human Factors Comprehensive Measurement System
Overall, the method consists of two parts: (i) estimation of mental load, and (ii) estimation of biomechanical training. The estimation of mental load first applies 3D geometric extraction of the working subject's infrastructure, including a computer vision processing stage for recovering the 3D information of the working subject using precise stereometric measurement equipment. The next step requires extracting attention from the concentrated interaction of the worker with the infrastructure. For this purpose, the worker is equipped with eye-tracking glasses and a set of wearable accelerometers, with sensor data transmitted wirelessly to a nearby PC for monitoring.
Data analysis will register self-centered video frames from eye-tracking glasses with visual information in the scene (e.g., artificial landmarks), thereby recovering the 3D geometric structure of gaze and visual cones directed towards the work subject. The final step is to link the semantics of the area of interest with task references, thus deriving an estimate of the task transition rate of the workers, leading to an assessment of mental workload.
These numbers will become part of a more complete cognitive spectrum in the future, which, in principle, includes psychological parameters (including emotions) that will allow us to assess a long-term general definition of human psychological states (including stress) based on a larger population and artificial studies.
In summary, these two approaches, psychological and biomechanical strain, together represent the main sources of psychophysiological stress for the work subject and will form the basis for future research in real work environments.
Development of digital twins for workstations
An important prerequisite for calculating mental load is defining the workspace, which represents the regions of interest (ROI) in the workstation environment that workers interact with during specific tasks. For example, during the assembly process of a specific part of a product, interactions apply to specific parts of the environment, such as the table. The basic principle for calculating concentration is that workers focus their attention when entering these environments, occupied by objects of interest not yet considered in this work, which will be completed in later work. However, 3D models are used not only to define task interaction areas but also to define the overall visualization of the work environment, as well as the overlay of interactions and related attention behaviors.

Figure 2. Development of a digital twin for the workstation environment, deriving eye movements related to the environment through the application of eye-tracking glasses, image recognition, and matching methods.
The primary task to be performed is to extract 3D models from real work environments (Figure 2). For this, we applied a mobile, practical stereo system that combines high precision, robust stereo projection texture, and efficient volume integration features, capable of easily capturing accurate 3D models of indoor scenes. Its method optimizes the stereo method of random point projection patterns and provides complete and robust results. The hardware is enclosed in a box containing three Basler dart cameras (2 monochrome cameras for stereo, 1 for RGB) and an active Kinect projector to apply reconstruction through occupancy grids and Iterative Closest Point (ICP) registration. Figure 3 shows typical results.
Estimation of mental load
To estimate the actual cognitive load from the interaction between workers and the environment, we decided to use eye-tracking glasses, which can continuously determine eye movements over time. In a previous work, we developed a method to estimate attention and mental load based on data from eye-tracking glasses in a factory-like laboratory environment. The eye-tracking glasses produce a continuous video stream of about 30Hz and eye gaze within video frames.
To locate the gaze extracted from the 3D environment system, we must first locate the environmental video, then assess the gaze rays and their geometric recovery relative to the 3D environment, and determine which ROI they belong to. To match the video frames with the 3D model, we placed ArUco markers (OpenCV toolbox) in the workstation where we expect the workers to gaze. Another method is to perform the matching process based not on reference markers but directly on video-based measurements from the scene.
•Relative to the work environment, over time (60 Hz), the gaze behavior of the visual axes of both eyes,
•Indicative estimates of mental stress[11],
•concentration on tasks[8], indicative estimates of sustained attention,
•perceived situation by estimating gaze on significant events[12],
•task transitions, indicating the load of switching attention between mental models[8](executive functions).
Assessment of biomechanical strain
To estimate biomechanical strain, we first analyzed human posture and summarized the overall body strain. We applied non-invasive wearable devices to workers to estimate relative positions and further derive the overall worker posture in space, ultimately receiving skeleton-based analysis results for further processing.
The wearable device we chose is the Perception Neuron, which represents a small, adaptive, multifunctional, and affordable motion capture technology. The modular system is based on NEURON, which consists of a 3-axis gyroscope, a 3-axis accelerometer, and a 3-axis magnetometer (IMU). This system applies proprietary embedded data fusion, human dynamics, and physics engine algorithms to achieve smooth motion with minimal latency. The Perception Neuron 9-axis sensor unit outputs data at 60fps or 120fps. The data stream is directed to a hub and can be transmitted to a computer in three different ways: (1) via WIFI, (2) via USB, or (3) recorded on the motherboard using a built-in micro-SD slot. The modular system on the notebook ultimately applies skeleton-based model displays.
To analyze the skeleton-based model, we use the biomechanics software toolkit BoB (Body Biomechanics Analysis). BoB is a package in which digitized human musculoskeletal simulation models can be created.
During the import task, BoB reads motion, force, and skeleton files. The skeleton file is applied to construct a mechanism containing joints, which use joint angles defined in the motion file for joint movement. Forces in the force file are applied as external forces to the mechanism. Then, BoB performs inverse dynamics analysis on the mechanism to calculate the joint torques required for joint movements in the presence of external forces and mass models. Based on the skeleton-based model, BoB is able to provide personalized estimates of characteristic biomechanical strain features, such as muscle forces, joint torques, and joint contact forces.

(a)

(b)
Figure 3. Results of the 3D information recovery process. (a) Generation of a three-dimensional model with areas of interest. (b) 3D gaze recovery using the camera's focus (green node) and estimated visual cone (green pyramid) along with 3D infrastructure.

Figure 4: The data stream received based on motion, force, and skeleton information, ultimately producing muscle and joint contact forces through inverse dynamics and optimization processes.
Figure 4 describes the data flow of the received motion, force, and skeletal files, as well as the muscle and joint contact forces generated by the optimization process.
There is no unique solution for the muscle force distribution required to generate joint torques, as there are typically about 40 joint torques to satisfy over 600 muscles in the muscle file. Therefore, an optimization method using a cost function is employed to identify one of the infinite possible solutions. By default, BoB minimizes the sum of the squares of muscle activations, where muscle activation is defined as instantaneous muscle force divided by the muscle's maximum isometric force. This cost function has been shown to correspond to reduced fatigue in the body. Users can choose other cost functions.
Experimental Results
The first result was obtained from a work subject oriented towards assembly, where frequent picking processes were accompanied by changes in direction and posture, characterizing the interaction pattern. Graz Medical University granted ethical certification (No. 31-243 ex 18/19). A typical female worker aged 51 wore the device for 11 minutes, or 660 seconds (Figure 5). During this period, a total of 15,840 video frames were captured, of which 27% were referenced in a fully automated manner associated with the ROI of the work environment infrastructure. The remaining video frames were manually attributed to the visible ROI in self-centered video frames. Heuristic concentration metrics were then calculated based on gaze reference ROI data (Paletta et al. [8]). This resulted in an average concentration level (standard deviation) of M=3.79 (SD=1.39), with level values ranging from 1 to 5 (Figure 6). During the 11 minutes of work, 153 task switches were calculated, with an average (SD) of 4.64 (2.78) switches, and a maximum of about 10.0 within 20-second intervals (Figure 7). These results indicate a very high task switching rate compared to normally reported functions [14].
In the final step, we associated the muscle strength of the scheduled muscle groups with the "occupational strain categories" through muscle analysis assessments provided by occupational therapists [15]. Support vector machine neural networks (Vapnik, 1995) were ultimately trained to map muscle strength to occupational stress. The first neural network achieved an accuracy of 89% in predictions within a 1-second time window.

Figure 5. Synchronized video: (Top left) Original video monitoring of wearable device assembly work, (Top right) Skeletal estimation, (Bottom left) Gaze video, (Bottom right) Visualization of muscle strength (green, yellow, red indicating low, medium, high muscle tension) and correct posture.

Figure 6. Heuristic concentration levels calculated for picking tasks and task-independent issues based on Paletta et al. [8] (from 1 (low) to 5 (high) (center).

Figure 7. Task switching (calculated according to Paletta et al. [8]), reflecting high cognitive load in phase 2.
Discussion, Conclusion, and Future Work
This experiment demonstrates that a lightweight, unobtrusive wearable device composed of eye-tracking glasses and motion capture sensors can be used to derive mental load and biomechanical strain. The heuristic scores extracted from gaze data reflect the high concentration and high task switching rates of workers. From the extracted quantities, we can easily conclude that cognitive workload is excessive, for example, workers should not focus continuously for more than 20-30 minutes, otherwise fatigue will worsen, which in turn significantly increases risk factors in the workplace.
In addition, detailed biomechanical strain information was mapped to scores from occupational experts, largely reflecting the experts' assessments.
Future work will focus on the application of more complex eye movement features and studies on larger populations, as we have already introduced biosensing technology for more fundamental processing of psychophysiological analysis of human factors and ergonomics in manufacturing environments.

Figure 8. 3D gaze geometry recovered from eye-tracking glasses overlaid with gaze points (orange) in a 3D model of the work environment and self-centered video frames (left).

Figure 9: Work in a modern factory, representing the picking process in an assembly work unit (top left). This worker is wearing eye-tracking glasses, gloves, and limbs equipped with wearable devices such as gyroscopes/accelerometers/magnetometers, and carrying a backpack with a notebook. From the wearable devices, we calculated a skeletal representation of human posture (right). The force on each muscle is quantified based on an estimated model derived from the current posture and plotted over time (blue, bottom). Additionally, we rated the sequences of video recordings using occupation-based activity analysis, providing annotations similar to ground truth by occupational therapists, describing low (green), medium (yellow), and high (red) levels of challenges to muscle function [15].
Acknowledgments
This work was supported by the Styrian Future Fund ("Zukunftsfonds Steiermark", project INCLUDE, No. 1036) and the Austrian Federal Ministry for Climate Action, Environment, Energy, Mobility, Innovation and Technology (BMK) within the projects MMASSIST (No. 858623) and FLEXIFF (No. 861264). Special thanks to Mr. Wolfgang Silly (KNAPP AG) for his strong support of the INCLUDE project.
References:
[1] A. Steinfeld, T. Fong, D. Kaber, M. Lewis, J. Scholtz and M. Goodrich, "Common metrics for human-robot interaction," inProc.Human-robot interaction, ACM SIGCHI/SIGART, 2006.
[2] C.-M. Huang and B. Mutlu, "Anticipatory robot control for efficient human-robot collaboration," inProc. ACM/IEEE HRI 2016, 2016.
[3] K. Fischer, L. Jensen, F. Kirstein, S. Stabinger,Ö. Erkent, D. Shukla and J. Piater, "The Effects of Social Gaze in Human-Robot Collaborative Assembly," inTapus A., André E., Martin JC., Ferland F., Ammi M. (eds), Social Robotics. ICSR 2015, Lecture Notes in Computer Science, vol 9388. Springer, Cham. https://doi.org/10.1007/978-3-319-25554-5_21, 2015.
[4] H. Admoni, A. Dragan, S. Srinivasa and B. Scassellati, "Deliberate delays during robot-to-human handovers improve compliance with gaze communication.," inProceedings of the ACM/IEEE International Conference on Human-Robot Interaction (HRI) (pp. 49, 2014.
[5] K. Santner, G. Fritz, L. Paletta and H. Mayer, "Visual recovery of saliency maps from human attention in 3D environments," inProc. International Conference on Robotics and Automation (ICRA), pp. 4297-4303, 2013.
[6] M. Klopschitz, R. Perko, G. Lodron, G. Paar and H. Mayer, "Projected Texture Fusion," in10th International Symposium on
Image and Signal Processing and Analysis (ISPA 2017), Ljubljana, Slovenia, September 18-20, 2017,.
[7] R. Newcombe, S. Izadi, O. Hilliges, D. Molyneaux, D. Kim, A. Fitzgibbon, A. Davison, P. Kohli, J. Shotton and S. Hodges, "KinectFusion: Real-Time Dense Surface Mapping and Tracking," in10th IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2011), Oct. 2011..
[8] L. Paletta, M. Pszeida, H. Ganster, F. Fuhrmann, W. Weiss, S. Ladstätter, A. Dini, S. Murg, H. Mayer, I. Brijacak and B. Reiterer, "Gaze based Human Factors Measurements for the Evaluation of Intuitive Human-Robot Collaboration in Real-time," inProc. 24th IEEE Conference on Emerging Technologies and Factory Automation, ETFA 2019, Zaragoza, Spain, 2019.
[9] S. Garrido-Jurado, F. Madrid-Cuevas, R. Muñoz-Salinas and M. Marín-Jiménez, "Automatic generation and detection of highly reliable fiducial markers under occlusion,"Pattern Recognition,pp. 47, 6 (June 2014), 2280-2292. doi=10.1016/j.patcog.2014.01.0, June 2014.
[10] L. Paletta, K. Santner, G. Fritz, A. Hofmann, G. Lodron, G. Thallinger and H. Mayer, "FACTS - A Computer Vision System for 3D Recovery and Semantic Mapping of Human Factors," inProc. 9th International Conference on Computer Vision Systems, ICVS 2013, LNCS 7963, pp. 62-72, Springer-Verlag Berlin Heidelberg, Sankt Petersburg, Russia, July 16-18, 2013, 2013.
[11] L. Paletta, M. Pszeida, B. Nauschnegg, T. Haspl and R. Marton, "Stress Measurement in Multi-tasking Decision Processes Using Executive Functions Analysis," inAyaz H. (eds) AHFE 2019. Advances in Neuroergonomics and Cognitive Engineering. Advances in Intelligent Systems and Computing, Springer, 2019, pp. vol 953, pp. 344-356.
[12] A. Dini, C. Murko, L. Paletta, S. Yahyanejad, U. Augsdörfer and M. Hofbaur, "Measurement and Prediction of Situation Awareness in Human-Robot Interaction based on a Framework of Probabilistic Attention," inProc. IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, Canada, 2017.
[13] J. Shippen and B. May, "BoB– biomechanics in MATLAB," inProceedings of 11th International Conference BIOMDLORE, 2016.
[14] M. Steyvers, G. E. Hawkins, F. Karayanidis and S. D. Brown, "A large-scale analysis of task switching practice effects across the lifespan,"PNAS ,pp. 17735-17740, September 3, 2019 116 (36) .
[15] H. Thomas, Occupation-based activity analysis, Thorofare, NJ: SLACK Incorporated., 2015.
Biomechanical Analysis,Biomechanical strain,Wearable devices,Motion capture of human body movements,Sports Biomechanics,Human body simulation modeling,Human Factors Engineering