COMBINED CONVOLUTIONAL NEURAL NETWORK (CNN) AND KEYPOINT-BASED METHOD FOR RECOGNIZING FINGER INTERACTION INTENT STATES
PDF (English)

Słowa kluczowe

prosthetics
assistive technology
vision-based control
robotics
hand keypoint detection
convolutional neural networks
human-machine interaction

Jak cytować

Marisova, Y. (2025). COMBINED CONVOLUTIONAL NEURAL NETWORK (CNN) AND KEYPOINT-BASED METHOD FOR RECOGNIZING FINGER INTERACTION INTENT STATES. European Journal of Interdisciplinary Issues, 2(4), 140–148. https://doi.org/10.5281/zenodo.19546851

Abstrakt

This work addresses the challenge of developing intuitive and accessible control systems for upper-limb prostheses with particular emphasis on pediatric applications. Conventional approaches - such as myoelectric and mechanical control often suffer from limitations related to signal stability, usability or reliability. The proposed hybrid method introduces computer vision as an alternative command channel, enabling vision-based prosthetic control driven by the interpretation of user intention, without relying on wearable bioelectrical or tactile sensors. Intention states are inferred from observed hand interactions using a contact-based labeling logic and a combined CNN and keypoint-based framework. A hardware-software prototype was implemented using a microcontroller platform and a 3D-printed robotic finger, in which a vision-based AI model performs real-time intention-state classification and generates corresponding actuation commands. A preliminary cost and feasibility assessment indicates the potential suitability of the proposed approach for cost-sensitive assistive devices. Considering the increasing number of individuals injured as a result of russian military aggression, this solution provides a highly relevant and socially significant assistive mechanism for both civilians and military personnel. Additionally, this system holds significant potential for users with congenital limb malformations, facilitating a more active lifestyle via responsive robotic prosthetic fingers. Consequently, the proposed system represents a practical and scalable technology that should be made accessible to every individual in need.

https://doi.org/10.5281/zenodo.19546851
PDF (English)

Bibliografia

Bazarevsky, V., Kartynnik, Y., Vakunov, A., Raveendran, K., & Grundmann, M. (2020). BlazePose: On-device real-time body pose tracking. CVPR Workshops. https://doi.org/10.48550/arXiv.2006.10204

Ghazaei, G., Alameer, A., Degenaar, P., Morgan, G., & Nazarpour, K. (2017). Deep learning-based artificial vision for grasp classification in myoelectric hands. Journal of Neural Engineering, 14(3), 036025. https://doi.org/10.1088/1741-2552/aa6802

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.48550/arXiv.1512.03385

Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.48550/arXiv.1709.01507

Klotzbuecher, M., & Bruyninckx, H. (2012). Coordinating robotic tasks and systems with rFSM statecharts. Journal of Software Engineering for Robotics, 3(1), 28–56. https://aisberg.unibg.it/retrieve/e40f7b86-2bd9-afca-e053-6605fe0aeaf2/52-254-1-PB.pdf

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2017). ImageNet classification with deep convolutional neural networks. Communications of the ACM, 60(6), 84–90. https://doi.org/10.1145/3065386

Kurmankhojayev, G., et al. (2021). Vision-based human intention recognition: A survey. IEEE Sensors Journal, (79), 30509–30555 https://doi.org/10.1007/s11042-020-09004-3

LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient-based learning applied to document recognition. In Proceedings of the IEEE. https://doi.org/10.1109/5.726791

Li, H., & Wu, X.-J. (2019). DenseFuse: A fusion approach to infrared and visible images. IEEE Transactions on Image Processing, 28(5), 2614–2623. https://blog.csdn.net/Pineapple_Daisy/article/details/136349283

Lin, M., Chen, Q., & Yan, S. (2014). Network in Network. In International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1312.4400

Marković, M., Dosen, S., Cipriani, C., Popović, D., & Farina, D. (2014). Stereovision and augmented reality for closed-loop control of grasping in hand prostheses. Journal of Neural Engineering, 11(4), 046001. https://doi.org/10.1088/1741-2560/11/4/046001

Molchanov, S., Gupta, S., Kim, K., & Kautz, J. (2015). Hand gesture recognition with 3D CNNs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. https://www.scribd.com/document/876243377/IEEE-Conference-Template-1

Mouchoux, J., Dosen, S., et al. (2022). Continuous semi-autonomous prosthesis control using a depth sensor on the hand. Frontiers in Neurorobotics, (15). https://doi.org/10.3389/fnbot.2022.814973

Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L.-C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.48550/arXiv.1801.04381

Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1409.1556

Srivastava, N., & Salakhutdinov, R. (2012). Multimodal learning with deep Boltzmann machines. In Advances in Neural Information Processing Systems. https://proceedings.neurips.cc/paper/2012/hash/af21d0c97db2e27e13572cbf59eb343d-Abstract.html

Tan, M., & Le, Q. (2019). EfficientNet. In International Conference on Machine Learning. https://doi.org/10.48550/arXiv.1905.11946

Villani, V., Pini, F., Leali, F., & Secchi, C. (2018). Survey on human–robot collaboration in industrial settings: Safety, intuitive interfaces and applications. Mechatronics, (55), 248–266. https://doi.org/10.1016/j.mechatronics.2018.02.009

Weiss, K., Khoshgoftaar, T., & Wang, D. (2016). A survey of transfer learning. Journal of Big Data, (3), 9. https://doi.org/10.1186/s40537-016-0043-6

Yosinski, J., Clune, J., Bengio, Y., & Lipson, H. (2014). How transferable are features in deep neural networks? In Advances in Neural Information Processing Systems. https://papers.nips.cc/paper_files/paper/2014/hash/532a2f85b6977104bc93f8580abbb330-Abstract.html

Zhang, F., Bazarevsky, V., Vakunov, A., Tkachenka, A., Sung, G., Chang, C.-L., & Grundmann, M. (2020). MediaPipe Hands: Real-time hand tracking. In CVPR Workshops. https://ijrpr.com/uploads/V6ISSUE11/IJRPR56027.pdf

Creative Commons License

Utwór dostępny jest na licencji Creative Commons Uznanie autorstwa – Użycie niekomercyjne 4.0 Międzynarodowe.

Prawa autorskie (c) 2025 Yana Marisova