US2016078287A1PendingUtilityA1

Method and system of temporal segmentation for gesture analysis

Assignee: KONICA MINOLTA LAB USA INCPriority: Aug 29, 2014Filed: Aug 29, 2014Published: Mar 17, 2016
Est. expiryAug 29, 2034(~8.1 yrs left)· nominal 20-yr term from priority
G06T 2207/10021G06K 9/4604G06K 9/00711G06K 9/00342G06K 9/6267G06V 20/64G06V 40/23G06V 20/49
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method, system and non-transitory computer readable medium for recognizing gestures are disclosed, the method includes capturing at least one three-dimensional (3D) video stream of data on a subject; extracting a time-series of skeletal data from the at least one 3D video stream of data; isolating a plurality of points of abrupt content change called temporal cuts, the plurality of temporal cuts defining a set of non-overlapping adjacent segments partitioning the time-series of skeletal data; identifying among the plurality of temporal cuts, temporal cuts of the time-series of skeletal data having a positive acceleration; and classifying each of the one or more pair of consecutive cuts with the positive acceleration as a gesture boundary.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for recognizing gestures, comprising:
 capturing at least one three-dimensional (3D) video stream of data on a subject;   extracting a time-series of skeletal data from the at least one 3D video stream of data;   isolating a plurality of points of abrupt content change and identifying each of the plurality of points of abrupt content change as a temporal cut, and wherein a plurality of temporal cuts define a set of non-overlapping adjacent segments partitioning the time-series of skeletal data;   identifying among the plurality of temporal cuts, temporal cuts of the time-series of skeletal data having a positive acceleration;   classifying each of the one or more pair of consecutive cuts with the positive acceleration as a gesture boundary.   
     
     
         2 . The method of  claim 1 , comprising:
 computing an estimated Maximum Mean Discrepancy (MMD) within the time-series of skeletal data; and   generating estimated temporal cuts among the time-series of skeletal data based on the estimated MMD.   
     
     
         3 . The method of  claim 2 , comprising:
 refining each of the estimated temporal cuts computed using the estimated MMD to generate a maximum rate of change of acceleration.   
     
     
         4 . The method of  claim 3 , comprising:
 generating the maximum rate of change of acceleration using a value of a hands-up decision function, wherein the hands-up decision function is a sum of vertical position of a left-hand joint and a right-hand joint at a time (t);   classifying a positive hands-up decision function a gesture; and   classifying a negative hands-up decision function as a non-gesture.   
     
     
         5 . The method of  claim 1 , comprising:
 classifying a positive rate of acceleration within a temporal cut as a beginning of the gesture; and   classifying a negative rate of acceleration within the temporal cut as an end of the gesture.   
     
     
         6 . The method of  claim 1 , comprising:
 inputting the time-series of skeletal data from the at least one 3D video stream of data and the gesture boundaries into a gesture recognition module; and   recognizing the gesture boundary as a type of gesture.   
     
     
         7 . A system for recognizing gestures, comprising:
 a video camera for capturing at least one three-dimensional (3D) video stream of data on a subject;   a module for extracting a time-series of skeletal data from the at least one 3D video stream of data; and   a processor configured to:
 isolate a plurality of points of abrupt content change and identifying each of the plurality of points of abrupt content change as a temporal cut, and wherein a plurality of temporal cuts define a set of non-overlapping adjacent segments partitioning the time-series of skeletal data; 
 identifying among the plurality of temporal cuts, temporal cuts of the time-series of skeletal data having a positive acceleration; 
 classifying each of the one or more pair of consecutive cuts with the positive acceleration as a gesture boundary. 
   
     
     
         8 . The system of  claim 7 , comprising:
 a display for displaying results generated by the processor in which one or more gesture boundaries from the time-series of skeletal data in a visual format.   
     
     
         9 . The system of  claim 7 , wherein the processor is configured to:
 compute an estimated Maximum Mean Discrepancy (MMD) within the time-series of skeletal data; and   generate estimated temporal cuts among the time-series of skeletal data based on the estimated MMD.   
     
     
         10 . The system of  claim 9 , wherein the processor is configured to:
 refine each of the estimated temporal cuts computed using the estimated MMD to generate a maximum rate of change of acceleration.   generate the maximum rate of change of acceleration using a value of a hands-up decision function, wherein the hands-up decision function is a sum of vertical position of a left-hand joint and a right-hand joint at a time (t);   classifying a positive hands-up decision function a gesture; and   classifying a negative hands-up decision function as a non-gesture.   
     
     
         11 . The system of  claim 10 , wherein the processor is configured to:
 classify a positive rate of acceleration within a temporal cut as a beginning of the gesture; and   classify a negative rate of acceleration within the temporal cut as an end of the gesture.   
     
     
         12 . The system of  claim 7 , comprising:
 a gesture recognition module configured to receive the time-series of skeletal data from the at least one 3D video stream of data and the gesture boundaries, and recognizing the gesture boundary as a type of gesture.   
     
     
         13 . The system of  claim 7 , wherein the video camera is a RGB-D camera, and wherein the RGB-D camera produces a time-series of RGB frames and depth frames. 
     
     
         14 . The system of  claim 7 , wherein the module for extracting a time-series of skeletal data from the at least one 3D video stream of data and the processor are in a standalone computer. 
     
     
         15 . A non-transitory computer readable medium containing a computer program storing computer readable code for recognizing gestures, the program being executable by a computer to cause the computer to perform a process comprising:
 capturing at least one three-dimensional (3D) video stream of data on a subject;   extracting a time-series of skeletal data from the at least one 3D video stream of data;   isolating a plurality of points of abrupt content change and identifying each of the plurality of points of abrupt content change as a temporal cut, and wherein a plurality of temporal cuts define a set of non-overlapping adjacent segments partitioning the time-series of skeletal data;   identifying among the plurality of temporal cuts, temporal cuts of the time-series of skeletal data having a positive acceleration;   classifying each of the one or more pair of consecutive cuts with the positive acceleration as a gesture boundary.   
     
     
         16 . The computer readable storage medium of  claim 15 , comprising:
 computing an estimated Maximum Mean Discrepancy (MMD) within the time-series of skeletal data; and   generating estimated temporal cuts among the time-series of skeletal data based on the estimated MMD.   
     
     
         17 . The computer readable storage medium of  claim 16 , comprising:
 refining each of the estimated temporal cuts computed using the estimated MMD to generate a maximum rate of change of acceleration.   
     
     
         18 . The computer readable storage medium of  claim 15 , comprising:
 generating the maximum rate of change of acceleration using a value of a hands-up decision function, wherein the hands-up decision function is a sum of vertical position of a left-hand joint and a right-hand joint at a time (t);   classifying a positive hands-up decision function a gesture; and   classifying a negative hands-up decision function as a non-gesture.   
     
     
         19 . The computer readable storage medium of  claim 15 , comprising:
 classifying a positive rate of acceleration within a temporal cut as a beginning of the gesture; and   classifying a negative rate of acceleration within the temporal cut as an end of the gesture.   
     
     
         20 . The computer readable storage medium of  claim 15 , comprising:
 inputting the time-series of skeletal data from the at least one 3D video stream of data and the gesture boundaries into a gesture recognition module; and   recognizing the gesture boundary as a type of gesture.

Join the waitlist — get patent alerts

Track US2016078287A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.