IP Library Patent Application 16266615
Patent Application
App. No. 16/266,615

METHODS AND APPARATUSES FOR PROCESSING VIDEO DATA

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/266,615
Abstract

A method and device for processing video data is provided. According to some embodiments, the method includes: recognizing at least one of a face or a piece of clothing from video data representing a scene; when the recognized face does not match a preset face or the recognized clothing does not match preset clothing, determining a user corresponding to the recognized face or recognized clothing to be a customer, the preset face or preset clothing corresponding to a greeter in the scene; performing a detection of at least one of facial expression, movement, or voice of the greeter from the video data, to generate a detection result; and determining a service quality of the greeter based on the detection result, to generate an assessment result.

Claims (65)

1 . A method, comprising:

recognizing at least one of a face or a piece of clothing from video data representing a scene;

when the recognized face does not match a preset face or the recognized clothing does not match preset clothing, determining a user corresponding to the recognized face or recognized clothing to be a customer, the preset face or preset clothing corresponding to a greeter in the scene;

performing a detection of at least one of facial expression, movement, or voice of the greeter from the video data, to generate a detection result; and

determining a service quality of the greeter based on the detection result, to generate an assessment result.

2 . The method of claim 1 , wherein performing the detection of at least one of facial expression, movement, or voice of the greeter from the video data, to generate the detection result, further comprises at least one of:

detecting and acquiring a facial expression of the greeter, matching the facial expression of the greeter against one or more preset facial expressions to obtain an expression matching result, and adding the expression matching result to the detection result;

detecting and acquiring a movement performed by the greeter, matching the movement of the greeter against one or more preset movements to obtain a movement matching result, and adding the movement matching result to the detection result; or

detecting and acquiring a voice transcript of the greeter, matching the voice transcript against one or more preset transcripts to obtain a voice matching result, and adding the voice matching result to the detection result.

3 . The method of claim 2 , wherein determining the service quality based on the detection result, to generate an assessment result, further comprises:

determining the service quality of the greeter based on at least one of the expression matching result, the movement matching result, or the voice matching result, and adding the service quality to the assessment result.

4 . The method of claim 1 , further comprising:

determining, based on the video data, whether the customer has left the scene; and

in response to the determination that the customer has left the scene, ending a monitoring session of the scene.

5 . The method of claim 4 , further comprising:

recording a start time and an end time for video data corresponding to the monitoring session, the start time being a point in time when a customer is determined in the scene, and the end time being a point in time when the customer is determined to have left the scene; and

linking the assessment result to the video data corresponding to the monitoring session.

6 . The method of claim 5 , further comprising:

determining a plurality of monitoring sessions;

performing a statistical analysis on a number of service sessions, service durations, and service qualities of the greeter based on video data corresponding to the plurality of monitoring sessions respectively and assessment results linked thereto, to generate a statistical result for the greeter; and

performing an attendance evaluation on the greeter based on the statistical result.

7 . A device for processing video data, comprising:

a memory storing instructions; and

a processor configured to execute the instructions to:

recognize at least one of a face or a piece of clothing from video data representing a scene;

when the recognized face does not match a preset face or the recognized clothing does not match preset clothing, determine a user corresponding to the recognized face or recognized clothing to be a customer, the preset face or preset clothing corresponding to a greeter in the scene;

perform a detection of at least one of facial expression, movement, or voice of the greeter from the video data, to generate a detection result; and

determine a service quality of the greeter based on the detection result, to generate an assessment result.

8 . The device of claim 7 , wherein the processor is further configured to execute the instructions to:

detect and acquire a facial expression of the greeter, match the facial expression of the greeter against one or more preset facial expressions to obtain an expression matching result, and add the expression matching result to the detection result;

detect and acquire a movement performed by the greeter, match the movement of the greeter against one or more preset movements to obtain a movement matching result, and add the movement matching result to the detection result; and

detect and acquire a voice transcript of the greeter, match the voice transcript against one or more preset transcripts to obtain a voice matching result, and add the voice matching result to the detection result.

9 . The device of claim 8 , wherein the processor is further configured to execute the instructions to:

determine the service quality of the greeter based on at least one of the expression matching result, the movement matching result, or the voice matching result, and add the service quality to the assessment result.

10 . The device of claim 7 , wherein the processor is further configured to execute the instructions to:

determine, based on the video data, whether the customer has left the scene; and

in response to the determination that the customer has left the scene, end a monitoring session of the scene.

11 . The device of claim 10 , wherein the processor is further configured to execute the instructions to

record a start time and an end time for video data corresponding to the monitoring session, the start time being a point in time when a customer is determined in the scene, and the end time being a point in time when the customer is determined to have left the scene; and

link the assessment result to the video data corresponding to the monitoring session.

12 . The device of claim 11 , wherein the processor is further configured to execute the instructions to:

determine a plurality of monitoring sessions;

perform a statistical analysis on a number of service sessions, service durations, and service qualities of the greeter based on video data corresponding to the plurality of monitoring sessions respectively and assessment results linked thereto, to generate a statistical result for the greeter; and

perform an attendance evaluation on the greeter based on the statistical result.

13 . A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to:

recognize at least one of a face or a piece of clothing from video data representing a scene;

when the recognized face does not match a preset face or the recognized clothing does not match preset clothing, determine a user corresponding to the recognized face or recognized clothing to be a customer, the preset face or preset clothing corresponding to a greeter in the scene;

perform a detection of at least one of facial expression, movement, or voice of the greeter from the video data, to generate a detection result; and

determine a service quality of the greeter based on the detection result, to generate an assessment result.

14 . The non-transitory computer-readable medium of claim 13 , wherein the instructions further cause the processor to perform at least one of:

detecting and acquiring a facial expression of the greeter, matching the facial expression of the greeter against one or more preset facial expressions to obtain an expression matching result, and adding the expression matching result to the detection result;

detecting and acquiring a movement performed by the greeter, matching the movement of the greeter against one or more preset movements to obtain a movement matching result, and adding the movement matching result to the detection result; or

detecting and acquiring a voice transcript of the greeter, matching the voice transcript of the greeter against one or more preset transcripts to obtain a voice matching result, and adding the voice matching result to the detection result.

15 . The non-transitory computer-readable medium of claim 14 , wherein the instructions further cause the processor to:

determine the service quality of the greeter based on at least one of the expression matching result, the movement matching result, or the voice matching result, and adding the service quality to the assessment result.

16 . The non-transitory computer-readable medium of claim 13 , wherein the instructions further cause the processor to:

determine, based on the video data, whether the customer has left the scene; and

in response to the determination that the customer has left the scene, end a monitoring session of the scene.

17 . The non-transitory computer-readable medium of claim 16 , wherein the instructions further cause the processor to:

record a start time and an end time for video data corresponding to the monitoring session, the start time being a point in time when a customer is determined in the scene, and the end time being a point in time when the customer is determined to have left the scene; and

link the assessment result to the video data corresponding to the monitoring session.

18 . The non-transitory computer-readable medium of claim 17 , wherein the instructions further cause the processor to:

determine a plurality of monitoring sessions;

perform a statistical analysis on a number of service sessions, service durations, and service qualities of the greeter based on video data corresponding to the plurality of monitoring sessions respectively and assessment results linked thereto, to generate a statistical result for the greeter; and

perform an attendance evaluation on the greeter based on the statistical result.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 5, 2024
From: EAST WEST BANK
To: KAMI VISION INCORPORATED
Reel/Frame 070792/0551 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Mar 28, 2022
From: KAMI VISION INCORPORATED
To: EAST WEST BANK
Reel/Frame 059512/0101 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2019
From: WANG, HUAYONG
To: SHANGHAI XIAOYI TECHNOLOGY CO., LTD.
Reel/Frame 048934/0611 →