IP Library › Granted Patent US 9,984,486
Granted Patent B2
US 9,984,486 · App. 15/064,362 · Granted May 29, 2018

Method and apparatus for voice information augmentation and displaying, picture categorization and retrieving

Inventor: Maochang Dang (Hangzhou, CN)
Assignee: Alibaba Group Holding Limited
G06T11/60G06F17/30026G06F17/30268G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,984,486
App. No.
15/064,362
Granted
May 29, 2018
Kind
B2
Abstract

A method of voice information augmentation including displaying a picture and identifying an object to be augmented in the picture. The method also includes receiving the voice information and establishing a mapping relationship between the voice information and the object to be augmented. The method accurately represents the content of the picture by augmenting the different objects in the picture with the different voice information.

Claims (88)

1. A computer-implemented method of voice information augmentation, the method comprising:

displaying a picture;

identifying a plurality of objects in the picture;

displaying a plurality of voice augmentation buttons at positions that lie laterally adjacent to the objects, the shapes of the buttons being substantially different than the shapes of the objects;

detecting a user input at a first voice augmentation button of the plurality of voice augmentation buttons, wherein the first voice augmentation button is associated with a first object of the plurality of objects;

receiving voice information; and

establishing and storing a mapping relationship between the voice information and the first object.

2. The method of claim 1 , further comprising extracting one or more keywords from the voice information.

3. The method of claim 2 , further comprising:

receiving a picture retrieval request that includes one or more requested keywords; and

identifying a stored picture having one or more keywords that match the one or more requested keywords.

4. The method of claim 1 , wherein identifying the plurality of objects in the picture further comprises:

recognizing an object in the picture by use of image structure segmentation and image non-structure segmentation; and

utilizing a special effect to display the object.

5. The method of claim 1 , wherein identifying the plurality of objects in the picture further comprises:

receiving a hand motioned delineation command, wherein the delineation command is utilized to identify a region;

based on the region and by use of image structure segmentation or image non-structure segmentation, determining an object positioned within the region; and

identifying the object positioned with the region as an object to be augmented.

6. The method of claim 1 , further comprising:

selecting the first object after the mapping relationship has been stored; and

displaying the voice information mapped to the first object.

7. A computer-implemented method of voice information augmentation, the method comprising:

displaying a picture;

identifying a plurality of objects in the picture;

displaying a plurality of voice augmentation buttons at positions corresponding to the objects;

detecting a user input at a first voice augmentation button of the plurality of voice augmentation buttons, wherein the first voice augmentation button is associated with a first object of the plurality of objects;

receiving voice information;

establishing and storing a mapping relationship between the voice information and the first object;

receiving an access permission configuring command; and

based on the command, configuring an access permission for the voice information corresponding to the first object, the access permission being selected from the group of public or private, and further comprising:

if a device performs the voice information augmentation, permitting the device to access and edit the voice information;

if the device does not perform the voice information augmentation and the access permission for the voice information is public, permitting the device to access the voice information but not to edit the voice information; and

if the device does not perform the voice information augmentation and the access permission for the voice information is private, not allowing the device to access or edit the voice information.

8. An apparatus for voice augmentation, the apparatus comprising:

a processor;

a first display unit to display a picture; and

a non-transitory computer-readable medium coupled to the processor, the non-transitory computer-readable medium having computer-readable instructions stored thereon to be executed by the processor, the instructions comprising:

identifying a plurality of objects in the picture;

receiving voice information;

displaying a plurality of voice augmentation buttons at positions that lie laterally adjacent to the plurality of objects, the shapes of the buttons being substantially different than the shapes of the objects;

detecting a user input at a first voice augmentation button of the plurality of voice augmentation buttons, wherein the first voice augmentation button is associated with a first object of the plurality of objects; and

establishing and storing a mapping relationship between the voice information and the first object.

9. The apparatus of claim 8 , wherein the instructions further comprise:

recognizing objects in the picture by use of image structure segmentation or image non-structure segmentation; and

displaying the objects with special effects.

10. The apparatus of claim 8 , wherein the instructions further comprise:

receiving a hand motioned delineation command,

determining a region based on the delineation command,

based on the region and by use of image structure segmentation and image non-structure segmentation, determining an object positioned within the region; and

identifying the object positioned with the region as an object to be augmented.

11. The apparatus of claim 8 , wherein the instructions further comprise:

receiving a picture retrieval request that includes one or more requested keywords; and

identifying a stored picture having one or more keywords in the voice information that match the one or more requested keywords such that a matching picture has all of the requested keywords.

12. An apparatus for voice augmentation, the apparatus comprising:

a processor;

a first display unit configured to display a picture; and

a non-transitory computer-readable medium coupled to the processor, the non-transitory computer-readable medium having computer-readable instructions stored thereon to be executed by the processor, the instructions comprising:

a first processing module configured to identify a plurality of objects in the picture;

a first voice information input module configured to receive voice information;

a first receiving module configured to receive a command for voice information augmentation, wherein the first display unit is configured to display a plurality of voice augmentation buttons at positions corresponding to the plurality of objects, wherein the first receiving module is configured to receive the command for voice information augmentation responsive to a first voice augmentation button of the plurality of voice augmentation buttons being user selected, wherein the first voice augmentation button is associated with a first object of the plurality of objects, and wherein the first processing module is further configured to establish and store a mapping relationship between the voice information and the first object;

an access permission configuration module configured to receive an access permission configuring command and to configure access permissions for the voice information mapped to the objects to be augmented based on the access permission configuring command, the access permission being selected from the group of public or private, and wherein, if the apparatus performs the voice information augmentation, the apparatus is permitted to access and edit the voice information, if another apparatus performs the voice information augmentation and the access permission for the voice information is public, the apparatus is permitted to access the voice information but not to edit the voice information, if another apparatus performs the voice information augmentation and the access permission for the voice information is private, the apparatus is not permitted to access or edit the voice information.

13. A non-transitory computer-readable storage medium having embedded therein program instructions, when executed by one or more processors of a device, causes the device to execute a process for voice information augmentation, the process comprising:

displaying a picture;

identifying a plurality of objects in the picture;

displaying a plurality of voice augmentation buttons at positions that lie laterally adjacent to the objects, the shapes of the buttons being substantially different than the shapes of the objects;

detecting a user input at a first voice augmentation button of the plurality of voice augmentation buttons, wherein the first voice augmentation button is associated with a first object of the plurality of objects;

receiving voice information; and

establishing a mapping relationship between the voice information and the first object.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the process further comprises extracting one or more keywords from the voice information.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the process further comprises:

receiving a picture retrieval request that includes one or more requested keywords; and

identifying a stored picture having one or more keywords that match the one or more requested keywords.

16. The non-transitory computer-readable storage medium of claim 13 , wherein identifying the plurality of objects in the picture comprises:

recognizing an object in the picture by use of image structure segmentation and image non-structure segmentation; and

utilizing a special effect to display the object.

17. The non-transitory computer-readable storage medium of claim 13 , wherein the process further comprises:

selecting the first object after the mapping relationship has been stored; and

displaying the voice information mapped to the first object.

18. A non-transitory computer-readable storage medium having embedded therein program instructions, when executed by one or more processors of a device, causes the device to execute a process for voice information augmentation, the process comprising:

displaying a picture;

identifying a plurality of objects in the picture;

displaying a plurality of voice augmentation buttons at positions corresponding to the objects;

detecting a user input at a first voice augmentation button of the plurality of voice augmentation buttons, wherein the first voice augmentation button is associated with a first object of the plurality of objects;

receiving voice information;

establishing a mapping relationship between the voice information and the first object;

receiving an access permission configuring command; and

based on the command, configuring an access permission for the voice information corresponding to the first object, the access permission being selected from the group of public or private, and further comprising:

if the device performs the voice information augmentation, permitting the device to access and edit the voice information; if the device does not perform the voice information augmentation and the access permission for the voice information is public, permitting the device to access the voice information but not to edit the voice information; and if the device does not perform the voice information augmentation and the access permission for the voice information is private, not allowing the device to access or edit the voice information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2016
From: DANG, MAOCHANG
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 038040/0562 →
Priority Claims (1)
CN 2015 1 0104464 · Mar 10, 2015 · national
Continuity (1)
Related Publication 20160267921A1 · Sep 15, 2016