IP Library › Granted Patent US 12,443,272
Granted Patent B2
US 12,443,272 · App. 17/689,460 · Granted Oct 14, 2025

Proactive actions based on audio and body movement

Inventors: Brian W. Temple (Santa Clara, CA); Devin W. Chalmers (Oakland, CA); Thomas G. Salter (Foster City, CA)
Assignee: Apple Inc.
G06F3/012G06F3/013G06F3/0481G06F3/16G06V40/174G06V40/20G10L25/51G10L25/78G10H2210/076
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,272
App. No.
17/689,460
Granted
Oct 14, 2025
Kind
B2
Abstract

Various implementations disclosed herein include devices, systems, and methods that determine that a user is interested in audio content by determining that a movement (e.g., a user's head bob) has a time-based relationship with detected audio content (e.g., the beat of music playing in the background). Some implementations involve obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to audio in the physical environment and the second sensor data corresponding to a body movement in the physical environment. A time-based relationship between one or more elements of the audio and one or more aspects of the body movement is identified based on the first sensor data and the second sensor data. An interest in content of the audio is identified based on identifying the time-based relationship. Various actions may be performed proactively based on identifying the interest in the content.

Claims (68)

1. A method comprising,

at an electronic device having a processor:

obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to audio in the physical environment and the second sensor data corresponding to a body movement in the physical environment;

selectively switching to different power states on the electronic device based on multiple triggers identified at the electronic device, where selectively switching through the different power states comprises:

detecting, based on the second sensor data, that the body movement corresponds to a type of body movement indicative of user interest in music;

based on detecting that the body movement corresponds to the type of body movement indicative of user interest in music, selectively triggering performance of audio analysis to determine if music is playing based on the first sensor data;

determining that music is playing based on the audio analysis;

based on determining that music is playing, selectively triggering performance of a comparison of audio elements with one or more aspects of the body movement, based on the first sensor data and the second sensor data;

identifying a time-based relationship between one or more elements of the audio and one or more aspects of the body movement via at least the comparison;

based on identifying the time-based relationship, selectively triggering performance of a source attribute identification process; and

identifying an interest in content of the audio based at least on the source attribute identification process;

determining to wait to provide one or more features based on a user state corresponding to the user being busy; and

providing one or more features based on identifying the interest in the content, wherein a timing of providing the one or more features comprises waiting based on the user state corresponding to the user being busy.

2. The method of claim 1 further comprising, based on identifying the interest in the content, presenting an identification of the content on the electronic device.

3. The method of claim 1 further comprising, based on identifying the interest in the content, presenting text corresponding to words in the content.

4. The method of claim 1 further comprising, based on identifying the interest in the content presenting a selectable option for:

replaying the content;

continuing to experience the content after leaving the physical environment;

purchasing the content;

downloading the content; or

adding the content to a playlist.

5. The method of claim 1 further comprising, based on identifying the interest in the content, identifying a characteristic of the content and identifying additional content based on the identified characteristic.

6. The method of claim 1 further comprising determining to limit providing features associated with the content based on determining the user state.

7. The method of claim 1 , wherein identifying the interest in the content is further based on determining that a voice in the physical environment is singing along with the content.

8. The method of claim 1 , wherein identifying the interest in the content is further based on an identified gaze direction corresponding to a device that is producing the audio or a direction relative to a head of the user corresponding to mental state indicative of user interest.

9. The method of claim 1 , wherein identifying the interest in the content is further based on an identified facial expression.

10. The method of claim 1 , wherein identifying the time-based relationship further comprises determining that a timing of a repeating body motion of the body movement matches a timing of a beat or rhythm of the content.

11. The method of claim 1 , wherein identifying the time-based relationship further comprises determining that lip movement of the body movement matches words of the content and identifying the interest in the content of the audio is based on identifying from the time-based relationship that the user is singing along or lip syncing to the words of the content.

12. The method of claim 1 , wherein identifying the time-based relationship further comprises determining that the body movement is a reaction to an event in the content.

13. The method of claim 1 , wherein the second sensor data corresponding to the body movement comprises:

image sensor data in which a portion of a body moves over time; or

motion sensor data from a motion sensor attached to a portion of the body.

14. The method of claim 1 further comprising identifying that multiple persons are interested in the content of the audio based on identifying time-based relationships using audio and body movement sensor data from the multiple persons in the physical environment.

15. A system comprising:

a non-transitory computer-readable storage medium; and

one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising:

obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to audio in the physical environment and the second sensor data corresponding to a body movement in the physical environment;

selectively switching to different power states on an electronic device based on multiple triggers identified at the electronic device, where selectively switching through the different power states comprises:

detecting, based on the second sensor data, that the body movement corresponds to a type of body movement indicative of user interest in music;

based on detecting that the body movement corresponds to the type of body movement indicative of user interest in music, selectively triggering performance of audio analysis to determine if music is playing based on the first sensor data;

determining that music is playing based on the audio analysis;

based on determining that music is playing, selectively triggering performance of a comparison of audio elements with one or more aspects of the body movement, based on the first sensor data and the second sensor data;

identifying a time-based relationship between one or more elements of the audio and one or more aspects of the body movement via at least the comparison;

based on identifying the time-based relationship, selectively triggering performance of a source attribute identification process; and

identifying an interest in content of the audio based at least on the source attribute identification process;

determining to wait to provide one or more features based on a user state corresponding to the user being busy; and

providing one or more features based on identifying the interest in the content, wherein a timing of providing the one or more features comprises waiting based on the user state corresponding to the user being busy.

16. The system of claim 15 , wherein the operations further comprise, based on identifying the interest in the content:

presenting an identification of the content on an electronic device;

presenting text corresponding to words in the content;

replaying the content;

continuing to experience the content after leaving the physical environment;

purchasing the content;

downloading the content; or

adding the content to a playlist.

17. The system of claim 15 , wherein identifying the interest in the content is further based on an identified gaze direction detected using images obtained via an image sensor.

18. A non-transitory computer-readable storage medium storing program instructions executable via one or more processors to perform operations comprising:

obtaining first sensor data and second sensor data corresponding to a physical environment, the first sensor data corresponding to audio in the physical environment and the second sensor data corresponding to a body movement in the physical environment;

selectively switching to different power states on an electronic device based on multiple triggers identified at the electronic device, where selectively switching through the different power states comprises:

detecting, based on the second sensor data, that the body movement corresponds to a type of body movement indicative of user interest in music;

based on detecting that the body movement corresponds to the type of body movement indicative of user interest in music, selectively triggering performance of audio analysis to determine if music is playing based on the first sensor data;

determining that music is playing based on the audio analysis;

based on determining that music is playing, selectively triggering performance of a comparison of audio elements with one or more aspects of the body movement, based on the first sensor data and the second sensor data;

identifying a time-based relationship between one or more elements of the audio and one or more aspects of the body movement via at least the comparison;

based on identifying the time-based relationship, selectively triggering performance of a source attribute identification process; and

identifying an interest in content of the audio based at least on the source attribute identification process;

determining to wait to provide one or more features based on a user state corresponding to the user being busy; and

providing one or more features based on identifying the interest in the content, wherein a timing of providing the one or more features comprises waiting based on the user state corresponding to the user being busy.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: TEMPLE, BRIAN W.; CHALMERS, DEVIN W.; SALTER, THOMAS G.
To: APPLE INC.
Reel/Frame 059198/0640 →
Continuity (2)
Provisional Application 63159503 · Mar 11, 2021
Related Publication 20220291743A1 · Sep 15, 2022
References Cited (24)
US 7674967B2 · Makino · 2010 [cited by examiner]
US 7806759B2 · McHale · 2010 [cited by examiner]
US 9202520B1 · Tang · 2015 [cited by examiner]
US 9286287B1 · Tierney · 2016 [cited by examiner]
US 10346472B2 · Alexandersson · 2019 [cited by examiner]
US 10368333B2 · Aggarwal · 2019 [cited by examiner]
US 10482124B2 · Chong et al. · 2019 [cited by applicant]
US 10846517B1 · Bulusu · 2020 [cited by examiner]
US 10860645B2 · Oh · 2020 [cited by examiner]
US 10897647B1 · Hunter Crawley · 2021 [cited by examiner]
US 11019300B1 · Baxendale · 2021 [cited by examiner]
US 11064266B2 · Trollope · 2021 [cited by examiner]
US 11561621B2 · Tang · 2023 [cited by examiner]
US 20120124604A1 · Small · 2012 [cited by examiner]
US 20140026156A1 · Deephanphongs · 2014 [cited by applicant]
US 20190042647A1 · Oh · 2019 [cited by examiner]
US 20190246936A1 · Garten et al. · 2019 [cited by applicant]
US 20200064458A1 · Giusti · 2020 [cited by examiner]
US 20200103967A1 · Bar-Zeev · 2020 [cited by applicant]
US 20200296521A1 · Wexler · 2020 [cited by examiner]
US 20210142792A1 · Wantland · 2021 [cited by examiner]
US 20220093101A1 · Krishnan · 2022 [cited by examiner]
US 20220291743A1 · Temple · 2022 [cited by examiner]
CN 107111642A · 2017 [cited by applicant]