IP Library › Granted Patent US 12,206,991
Granted Patent B2
US 12,206,991 · App. 17/711,923 · Granted Jan 21, 2025

Body language detection and microphone control

Inventors: Robert Michael Jordan (Orlando, FL); Howard Mall (Winter Springs, FL)
Assignee: Universal City Studios LLC
H04N23/695G06V40/176G06V40/23G06V40/28H04R1/08H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,206,991
App. No.
17/711,923
Granted
Jan 21, 2025
Kind
B2
Abstract

A system includes a gimbal, a shotgun microphone coupled to the gimbal, a camera, and at least one processor. The at least one processor is configured to receive data indicative of an image or video feed from the camera. The at least one processor is also configured to determine, based on the data indicative of the image or video feed, a primary human speaker among a group of humans and a location of the primary human speaker. The at least one processor is also configured to control the gimbal to point the shotgun microphone at the location of the primary human speaker.

Claims (54)

1. A system, comprising:

a gimbal;

a shotgun microphone coupled to the gimbal;

a camera; and

at least one processor configured to:

receive first data indicative of an image or video feed from the camera;

determine, based on the first data indicative of the image or video feed, a primary human speaker among a group of humans and a location of the primary human speaker;

control the gimbal to point the shotgun microphone at the location of the primary human speaker after determining the location of the primary human speaker; and then

receive, via the shotgun microphone, second data indicative of a sound captured by the shotgun microphone;

determine, based on the second data indicative of the sound captured by the shotgun microphone, a command uttered by the primary human speaker; and

execute the command by controlling a show element of the system.

2. The system of claim 1 , comprising an electronic screen configured to output a digital avatar, wherein the at least one processor is configured to:

determine a characteristic of the digital avatar based on an identity of the primary human speaker, the location of the primary human speaker, or both; and

control the electronic screen to output the digital avatar having the characteristic.

3. The system of claim 1 , comprising the show element, wherein the show element is actuatable between a plurality of show element positions, and wherein the at least one processor is configured to:

determine a show element position of the plurality of show element positions based on an identity of the primary human speaker, the location of the primary human speaker, or both; and

execute the command by controlling the show element to actuate the show element to the show element position.

4. The system of claim 1 , wherein the at least one processor is configured to execute a body language detection algorithm to identify, based on the first data indicative of the image or video feed, a hand gesture indicative of the primary human speaker.

5. The system of claim 1 , wherein the at least one processor is configured to execute a body language detection algorithm to identify, based on the first data indicative of the image or video feed, a facial expression indicative of the primary human speaker.

6. The system of claim 1 , wherein the at least one processor is configured to execute the command by controlling a characteristic of a digital avatar displayed on an electronic screen corresponding to the show element of the system.

7. The system of claim 1 , wherein the show element comprises an electronic screen, a physical show prop, a light, or any combination thereof.

8. The system of claim 1 , wherein the show element comprises an electronic screen, and the at least one processor is configured to execute the command by controlling the electronic screen to change a color of a digital avatar presented on the electronic screen.

9. The system of claim 1 , wherein the show element comprises a physical show prop, and the at least one processor is configured to execute the command by controlling the physical show prop to change a position of the physical show prop.

10. A system, comprising:

a microphone assembly;

a camera; and

at least one processor configured to:

receive first data indicative of an image or video feed from the camera;

determine, based on the first data indicative of the image or video feed, a primary human speaker among a group of humans and a location of the primary human speaker;

control the microphone assembly to point a microphone of the microphone assembly toward the location of the primary human speaker after determining the location of the primary human speaker; and then

receive, via the microphone, second data indicative of a sound captured by the microphone;

determine, based on the second data indicative of the sound captured by the microphone, a command uttered by the primary human speaker; and

execute the command by controlling a show element of the system.

11. The system of claim 10 , wherein the microphone assembly comprises a gimbal, and the at least one processor is configured to control the microphone assembly to point the microphone toward the location of the primary human speaker by controlling the gimbal.

12. The system of claim 11 , wherein the gimbal comprises a motor and the at least one processor is configured to control the microphone assembly to point the microphone toward the location of the primary human speaker by controlling the motor of the gimbal.

13. The system of claim 10 , wherein the microphone assembly comprises an array of microphones, and the at least one processor is configured to:

identify, from the array of microphones, a first microphone that does not correspond to the location of the primary human speaker; and

deactivate the first microphone.

14. The system of claim 13 , wherein the at least one processor is configured to:

identify, from the array of microphones, a second microphone that corresponds to the location of the primary human speaker; and

receive, from the second microphone, third data indicative of additional sound captured by the second microphone.

15. The system of claim 10 , comprising a speaker, wherein the at least one processor is configured to control the speaker to output audio corresponding to the sound captured by the microphone.

16. The system of claim 10 , comprising an electronic screen, wherein the at least one processor is configured to control a visual presentation displayed on the electronic screen based on the second data.

17. One or more tangible, non-transitory, computer readable media comprising instructions thereon that, when executed by at least one processor, cause the at least one processor to:

receive, from a camera, first data indicative of an image or video feed capturing a group of humans;

determine, via a body language detection algorithm that receives the first data indicative of the image or video feed, a primary human speaker and a location of the primary human speaker;

control a motorized gimbal to point a shotgun microphone at the location of the primary human speaker;

receive, via the shotgun microphone, second data indicative of a sound captured by the shotgun microphone;

determine, based on the second data indicative of the sound captured by the shotgun microphone, a command uttered by the primary human speaker; and

control a show element based on the command.

18. The one or more tangible, non-transitory, computer readable media of claim 17 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to control the show element or an additional show element based on an identity of the primary human speaker, the location of the primary human speaker, or both.

19. The one or more tangible, non-transitory, computer readable media of claim 18 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to:

control a first characteristic of the show element based on the identity of the primary human speaker, the location of the primary human speaker, or both; and

control a second characteristic of the show element based on the command such that the show element includes the first characteristic and the second characteristic simultaneously.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2022
From: JORDAN, ROBERT MICHAEL; MALL, HOWARD
To: UNIVERSAL CITY STUDIOS LLC
Reel/Frame 059486/0221 →
Continuity (1)
Related Publication 20230319416A1 · Oct 5, 2023
References Cited (20)
US 6219645B1 · Byers · 2001 [cited by examiner]
US 8150063B2 · Chen et al. · 2012 [cited by applicant]
US 8395653B2 · Feng et al. · 2013 [cited by applicant]
US 8787113B2 · Turbahn et al. · 2014 [cited by applicant]
US 9185361B2 · Curry · 2015 [cited by applicant]
US 9445193B2 · Tico et al. · 2016 [cited by applicant]
US 10074251B2 · Oh et al. · 2018 [cited by applicant]
US 10601385B2 · Moberg et al. · 2020 [cited by applicant]
US 11184560B1 · Mese · 2021 [cited by examiner]
US 11259112B1 · Marti et al. · 2022 [cited by applicant]
US 20110154266A1 · Friend et al. · 2011 [cited by applicant]
US 20150088515A1 · Beaumont · 2015 [cited by examiner]
US 20150167956A1 · Vaidya · 2015 [cited by examiner]
US 20160277863A1 · Cahill · 2016 [cited by examiner]
US 20170186441A1 · Wenus et al. · 2017 [cited by applicant]
US 20200192485A1 · VanBlon · 2020 [cited by examiner]
US 20210134293A1 · Zurek et al. · 2021 [cited by applicant]
US 20210256987A1 · Edlin · 2021 [cited by examiner]
US 20210279475A1 · Tusch · 2021 [cited by examiner]
PCT/US2023/017107 International Search Report and Written Opinion mailed Jun. 19, 2023. [cited by applicant]
Cited By (1)
US 12,326,968