Motion taxonomy for manipulation embedding and recognition
A method for motion recognition and embedding is disclosed. The method may include receiving a plurality of frames of an input video for extracting a feature vector of a motion in the plurality of frames, generating a plurality of sets of one or more motion component bits based on the feature vector and a plurality of classifiers, the plurality of sets corresponding to the plurality of classifiers, each set of one or more motion component bits representing a physical or mechanical attribute of the motion; and generating a motion code for a machine to execute the motion by combining the plurality of sets of one or more motion component bits. Other aspects, embodiments, and features are also claimed and described.
1. A method for motion recognition and embedding, comprising:
receiving a plurality of frames of an input video for extracting a feature vector of a motion in the plurality of frames;
generating a plurality of sets of one or more motion component bits based on the feature vector and a plurality of classifiers, the plurality of sets corresponding to the plurality of classifiers, a first set of one or more motion component bits representing a time-based physical or mechanical attribute of the motion;
generating a motion code corresponding to the motion of the input video based on the plurality of sets of one or more motion component bits; and
communicating the motion code to a machine to execute an action by combining the plurality of sets of one or more motion component bits.
2. The method of claim 1 , wherein the plurality of classifiers comprise an interaction type, an active trajectory type, and a passive trajectory type.
3. The method of claim 2 , wherein the interaction type indicates whether the motion is a contact motion,
wherein the contact motion comprises an engagement type and a contact duration type,
wherein the engagement type indicates whether the contact motion causes deformation on an object, and
wherein the contact duration type indicates whether the contact motion persists for a predetermined period of time.
4. The method of claim 2 , wherein the active trajectory type comprises a prismatic type, a revolute type, and a recurrence type,
wherein the prismatic type indicates whether the motion is along a single axis, a plane, or a manifold space,
wherein the revolute type indicates whether the motion is rotational, and
wherein the recurrence type indicates whether the motion is cyclical.
5. The method of claim 2 , wherein the passive trajectory type indicates whether the motion is with a passive object.
6. The method of claim 1 , wherein the generating the motion code comprises combining the plurality of sets of one or more motion component bits in an ordered sequence.
7. The method of claim 1 , wherein the motion code comprises a binary encoded string for a machine to execute the motion.
8. The method of claim 1 , further comprising:
selecting a highest probability value in a verb class of the motion; and
generating a probability distribution vector for the verb class of the motion based on the extracted feature vector and the highest probability value.
9. The method of claim 8 , further comprising:
combining the probability distribution vector and the motion code into a single feature vector.
10. The method of claim 9 , further comprising:
inserting the single feature vector into an artificial neural network (ANN); and
producing a final verb class probability distribution of the motion from the ANN based on the single feature vector.
11. The method of claim 1 , further comprising:
incorporating a semantic feature of an object into the motion into the feature vector; and
incorporating the semantic feature into the plurality of classifiers.
12. The method of claim 1 , wherein the machine performs an emulating motion corresponding to the motion of the input video, based upon the motion code.
13. The method of claim 1 , wherein the machine performs a consequential action as a result of the motion code.
14. An apparatus for motion recognition and embedding, comprising:
a processor; and
a memory communicatively coupled to the processor,
wherein the processor and the memory are configured to:
receive a plurality of frames of an input video for extracting a feature vector of a motion in the plurality of frames;
generate a plurality of sets of one or more motion component bits based on the feature vector and a plurality of classifiers, the plurality of sets corresponding to the plurality of classifiers, each set of one or more motion component bits representing a physical or mechanical attribute of the motion, wherein the plurality of classifiers comprise an interaction type, a prismatic type, a revolute type, a recurrence type, and a passive trajectory type, wherein the prismatic type indicates whether the motion is along a single axis, a plane, or a manifold space, wherein the revolute type indicates whether the motion is rotational, and wherein the recurrence type indicates whether the motion is cyclical; and
generate a motion code for a machine to execute the motion by combining the plurality of sets of one or more motion component bits.
15. The apparatus of claim 14 , wherein the generating the motion code comprises combining the plurality of sets of one or more motion component bits in an ordered sequence.
16. The apparatus of claim 14 , wherein the processor and the memory are further configured to:
select a highest probability value in a verb class of the motion; and
generate a probability distribution vector for the verb class of the motion based on the extracted feature vector and the highest probability value.
17. The apparatus of claim 16 , wherein the processor and the memory are further configured to:
combine the probability distribution vector and the motion code into a single feature vector.
18. The apparatus of claim 17 , wherein the processor and the memory are further configured to:
insert the single feature vector into an artificial neural network (ANN); and
produce a final verb class probability distribution of the motion from the ANN based on the single feature vector.
19. The apparatus of claim 14 , wherein the generating the motion code comprises generating the motion code based on an objective function (L M ):
L M =−Σ k=1 5 Σ l=1 C k λ k m l k log( f l k ( x )),
where λ k is a constant weight and m l k is a l th element of a ground truth on-hot vector for a kth set of the one or more motion component bits.
20. The apparatus of claim 14 , wherein the processor and the memory are further configured to:
incorporate a semantic feature of an object into the motion into the feature vector; and
incorporate the semantic feature into the plurality of classifiers.
21. A method for motion recognition and embedding, comprising:
receiving a plurality of frames of an input video for extracting a feature vector of a motion in the plurality of frames;
generating a plurality of sets of one or more motion component bits based on the feature vector and a plurality of classifiers, the plurality of sets corresponding to the plurality of classifiers, each set of one or more motion component bits representing a physical or mechanical attribute of the motion;
generating a motion code corresponding to the motion of the input video based on an objective function (L M ): L M =−Σ k=1 5 Σ l=1 C k λ k m l k log(f l k (x)), where λ k is a constant weight and m l k is a l th element of a ground truth on-hot vector for a kth set of the one or more motion component bits; and
communicating the motion code to a machine to execute an action by combining the plurality of sets of one or more motion component bits.