IP Library Granted Patent US 10,839,853
Granted Patent B2
US 10,839,853 · App. 16/167,564 · Granted Nov 17, 2020

Audio encoding for functional interactivity

Inventors: Viswanathan Iyer (Santa Clara, CA); Kartik Parija (Bangalore, IN)
Assignee: ADORI LABS, INC.
G11B20/12G06F3/0482G06F16/61G06F16/635G06F16/638G06F16/686G11B20/10527H04H20/31H04H60/33H04H60/74H04N21/4334H04N21/4398H04N21/8352G10L19/0018G11B2020/1265H04H2201/37H04H2201/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,839,853
App. No.
16/167,564
Granted
Nov 17, 2020
Kind
B2
Abstract

Some examples include receiving audio content through a microphone of an electronic device and determining whether embedded data is included in the received audio content. The electronic device may decode the received audio content to extract the embedded data. In addition, the electronic device may perform at least one of: sending a communication to a computing device over a network based on the extracted embedded data, or presenting information on a display of the electronic device based on the extracted embedded data.

Claims (70)

1. A system comprising:

an audio encoder for embedding data into audio content; and

a first computing device in communication with the audio encoder, the first computing device including a processor configured by executable instructions to perform operations comprising:

receiving audio content for distribution to a plurality of electronic devices;

presenting a user interface to enable a user to select data to be related to the audio content for distribution to the plurality of electronic devices, the user interface presenting a plurality of different data categories selectable in the user interface for selecting first data to at least one of associate with the audio content or embed in the audio content;

decoding a portion of the received audio content to detect a start of frame indicator to determine that the audio content already has second data embedded in the audio content;

based on detecting at least the start of frame indicator, extracting a psychoacoustic mask from the received audio content;

subtracting the psychoacoustic mask from the received audio content to remove the embedded second data;

receiving, via the user interface, a user selection providing an indication of the first data to at least one of associate with the audio content or embed in the audio content;

based at least in part on determining that the first data exceeds a threshold size, sending, by the first computing device, the first data to a second computing device to be provided for download to the electronic devices; and

based at least in part on determining that the first data exceeds the threshold size, sending, to the audio encoder, third data to embed in the audio content as embedded third data, the third data including information for identifying the first data following extraction of the embedded third data by individual electronic devices of the plurality of electronic devices.

2. The system as recited in claim 1 , wherein the third data includes a content identifier (ID), the operations further comprising:

generating, by the first computing device, a fingerprint from the audio content;

storing, by the first computing device, the fingerprint in association with the content ID; and

sending, by the first computing device, the first data, the content ID, and the fingerprint to the second computing device to provide the first data for storage in a content management system (CMS) repository associated with the second computing device,

wherein the plurality of electronic devices are able to access the first data in the CMS repository based at least in part on the embedded third data embedded in the audio content.

3. The system as recited in claim 1 , the operations further comprising receiving, from the second computing device, information about a plurality of users of the plurality of electronic devices, respectively, based on the second computing device receiving communications from the plurality of electronic devices based at least in part on the third data.

4. The system as recited in claim 3 , wherein the received information about the plurality of users includes at least one of location information, demographic information, listening duration information, or an action performed in response to at least one of the first data or the third data.

5. The system as recited in claim 1 , the operations further comprising sending the audio content with the embedded third data as at least one of: streaming content sent over a network, or broadcasted content sent as a radio wave.

6. The system as recited in claim 1 , the operations further comprising storing the audio content for subsequent access by one or more of the electronic devices, wherein the one or more electronic devices are subsequently able to access the first data at the second computing device based on accessing the stored audio content with the embedded third data, extracting the embedded third data from the audio content, using the third data to obtain a fingerprint of the audio content, and accessing the first data based on timing information determined at least in part from the fingerprint.

7. The system as recited in claim 1 , the operations further comprising:

prior to extracting the psychoacoustic mask from the received audio content, decoding the received audio content to extract at least a source identifier (ID) from the second data embedded in the received audio content;

comparing the source ID extracted from the second data embedded in the received audio content with a current source ID; and

based at least in part on the extracted source ID being different from the current source ID, performing the subtracting the psychoacoustic mask from the received audio content.

8. A method comprising:

receiving, by one or more processors, audio content for distribution to a plurality of electronic devices;

presenting a user interface to enable a user to select first data to be related to the audio content for distribution to the plurality of electronic devices;

decoding a portion of the received audio content to detect a start of frame indicator to determine that the audio content already has second data embedded in the audio content;

based on detecting at least the start of frame indicator, extracting a psychoacoustic mask from the received audio content;

subtracting the psychoacoustic mask from the received audio content to remove the embedded second data;

receiving, via the user interface, a user selection providing an indication of the first data to at least one of associate with the audio content or embed in the audio content;

sending, by the first computing device, the first data to a second computing device to provide for download to the electronic devices; and

sending, to an audio encoder, third data to embed in the audio content as embedded third data, the third data including information for identifying the first data following extraction of the embedded third data by individual electronic devices of the plurality of electronic devices.

9. The method as recited in claim 8 , wherein the third data includes a content identifier (ID), the method further comprising:

generating, by the first computing device, a fingerprint from the audio content;

storing, by the first computing device, the fingerprint in association with the content ID; and

sending, by the first computing device, the first data, the content ID, and the fingerprint to the second computing device to provide the first data for storage in a content management system (CMS) repository associated with the second computing device,

wherein the plurality of electronic devices are able to access the first data in the CMS repository based at least in part on the embedded third data embedded in the audio content.

10. The method as recited in claim 8 , wherein the user interface presents a plurality of virtual controls selectable for selecting one or more of a plurality of different data categories, respectively, for selecting the first data to at least one of associate with the audio content or embed in the audio content.

11. The method as recited in claim 8 , further comprising:

comparing a size of the first data with a threshold size for embedding data into the audio content; and

based at least in part on determining that the first data exceeds the threshold size, determining to not embed the first data into the audio content, sending the first data to the second computing device to be provided for download, and sending, to the audio encoder, the third data including information for identifying the first data.

12. The method as recited in claim 8 , further comprising sending the audio content with the embedded third data as at least one of: streaming content sent over a network, or broadcasted content sent as a radio wave.

13. The method as recited in claim 8 , further comprising storing the audio content for subsequent access by one or more of the electronic devices, wherein the one or more electronic devices are subsequently able to access the first data at the second computing device based on accessing the stored audio content with the embedded third data, extracting the embedded third data from the audio content, using the third data to obtain a fingerprint of the audio content, and accessing the first data based on timing information determined at least in part from the fingerprint.

14. The method as recited in claim 8 , further comprising:

prior to extracting the psychoacoustic mask from the received audio content, decoding the received audio content to extract at least a source identifier (ID) from the second data embedded in the received audio content;

comparing the source ID extracted from the second data embedded in the received audio content with a current source ID; and

based at least in part on the extracted source ID being different from the current source ID, performing the subtracting the psychoacoustic mask from the received audio content.

15. A computing device comprising:

one or more processors configured by executable instructions to perform operations comprising:

receiving audio content for distribution to a plurality of electronic devices;

presenting a user interface to enable a user to select first data to be related to the audio content for distribution to the plurality of electronic devices;

decoding a portion of the received audio content to detect a start of frame indicator to determine that the audio content already has second data embedded in the audio content;

based on detecting at least the start of frame indicator, extracting a psychoacoustic mask from the received audio content;

subtracting the psychoacoustic mask from the received audio content to remove the embedded second data;

receiving, via the user interface, a user selection providing an indication of the first data to at least one of associate with the audio content or embed in the audio content;

based at least in part on determining that the first data exceeds a threshold size, sending, by the first computing device, the first data to a second computing device to be provided for download to the electronic devices; and

based at least in part on determining that the first data exceeds the threshold size, sending, to an audio encoder, third data to embed in the audio content as embedded third data, the third data including information for identifying the first data following extraction of the embedded third data by individual electronic devices of the plurality of electronic devices.

16. The computing device as recited in claim 15 , wherein the third data includes a content identifier (ID), the operations further comprising:

generating, by the first computing device, a fingerprint from the audio content;

storing, by the first computing device, the fingerprint in association with the content ID; and

sending, by the first computing device, the first data, the content ID, and the fingerprint to the second computing device to provide the first data for storage in a content management system (CMS) repository associated with the second computing device,

wherein the plurality of electronic devices are able to access the first data in the CMS repository based at least in part on the embedded third data embedded in the audio content.

17. The computing device as recited in claim 15 , wherein the user interface presents a plurality of virtual controls selectable for selecting one or more of a plurality of different data categories, respectively, for selecting the first data to at least one of associate with the audio content or embed in the audio content.

18. The computing device as recited in claim 15 , the operations further comprising sending the audio content with the embedded third data as at least one of: streaming content sent over a network, or broadcasted content sent as a radio wave.

19. The computing device as recited in claim 15 , the operations further comprising storing the audio content for subsequent access by one or more of the electronic devices, wherein the one or more electronic devices are subsequently able to access the first data at the second computing device based on accessing the stored audio content with the embedded third data, extracting the embedded third data from the audio content, using the third data to obtain a fingerprint of the audio content, and accessing the first data based on timing information determined at least in part from the fingerprint.

20. The computing device as recited in claim 15 , the operations further comprising:

prior to extracting the psychoacoustic mask from the received audio content, decoding the received audio content to extract at least a source identifier (ID) from the second data embedded in the received audio content;

comparing the source ID extracted from the second data embedded in the received audio content with a current source ID; and

based at least in part on the extracted source ID being different from the current source ID, performing the subtracting the psychoacoustic mask from the received audio content.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2024
From: ADORI LABS, INC.
To: ADORI AI, INC.
Reel/Frame 069316/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2018
From: IYER, VISWANATHAN; PARIJA, KARTIK
To: ADORI LABS, INC.
Reel/Frame 047269/0014 →
Continuity (2)
Provisional Application 62576620 · Oct 24, 2017
Related Publication 20190122698A1 · Apr 25, 2019
Cited By (1)
US 12,354,591