IP Library Granted Patent US 12,238,497
Granted Patent B2
US 12,238,497 · App. 18/507,661 · Granted Feb 25, 2025

Systems, methods, apparatus, and computer-readable media for gestural manipulation of a sound field

Inventors: Pei Xiang (San Diego, CA); Erik Visser (San Diego, CA)
Assignee: QUALCOMM Incorporated
H04R5/04G06F3/017H04R3/005H04S7/303H04R2203/12H04R2430/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,238,497
App. No.
18/507,661
Granted
Feb 25, 2025
Kind
B2
Abstract

Gesture-responsive modification of a generated sound field is described.

Claims (34)

1. A device comprising:

a memory configured to store a representation of a gesture; and

one or more processors coupled to the memory, the one or more processors configured to:

recognize at least one movement of a user as the representation of the gesture, wherein the representation of the gesture is mapped to a command, wherein the command is context-dependent, wherein the command is produced in response to the representation of the gesture that is appropriate for the current context, wherein one gesture appropriate for the current context is to ignore the representation of the gesture to reduce volume when a system in already in a muted state;

apply the at least one recognized movement of the user to determine an indicated change to a sound field produced by an array of loudspeakers;

synthesize a modified sound field produced by the array of loudspeakers to implement the indicated change.

2. The device of claim 1 , wherein the representation of gesture is interpreted as one of a plurality of separate patterns, and decisions by the one or more processors are made to synthesize the corresponding sound field associated with the separate patterns.

3. The device of claim 1 , wherein the one or more processors are configured to extract one or more features used to detect and locate regions of interest.

4. The device of claim 3 , wherein the regions of interest include the user's eyes.

5. The device of claim 3 , wherein the regions of interest include the user's hands.

6. The device of claim 3 , wherein the regions of interest include the user's mouth.

7. The device of claim 3 , wherein the regions of interest include the user's body.

8. The device of claim 1 , wherein the one or more processors are configured to control the array of loudspeakers to generate beams in different directions to support gesture control independently for different users located in the different directions.

9. The device of claim 1 , wherein a voice command is used to enter a gesture control mode.

10. The device of claim 1 , wherein the at least one recognized movement of the user includes face recognition, voice recognition, or both face recognition and voice recognition for user identification or user location.

11. The device of claim 1 , wherein the command interpreter integrated into the one or more processors is configured to disable changes to a current sound field configuration or enable changes to the current sound field configuration.

12. The device of claim 1 , wherein the one or more processors are configured to recognize the at least one movement based on depth information.

13. The device of claim 12 , further comprising two or more cameras configured to generate the depth information.

14. The device of claim 12 , further comprising a projector configured to project a pattern of stripes, a pattern of dots, or both a pattern of stripes and dots onto a part of the user and estimate depths of surface points of the part of the user.

15. The device of claim 1 , wherein the at least one recognized movement of the user is based on an array of ultrasound transducers used configured to perform spatial imaging.

16. The device of claim 1 , further comprising an on-screen display to provide feedback for the gesture; wherein the feedback is a bar or a dial to display a change in beam intensity, beam direction, or dynamic range.

17. The device of claim 16 , wherein the on-screen display is configured to display the feedback for the gesture on a bar on the screen or dial on the screen to represent a change in beam intensity, beam direction, or dynamic range.

18. The device of claim 16 , wherein the on-screen display is configured to display the an error indication of an invalid gesture.

19. The device of claim 1 , wherein the at least one recognized movement of the user that represents the gesture is at least one among the following gestures: two-hand gesture, hand-and-head gesture, hand and body gesture, and hand to ear gesture.

20. The device of claim 1 , wherein the at least one recognized movement of the user that represents the gesture is at least one among the following gestures: a clockwise hand movement, a counterclockwise hand movement, and a hand rotation, hand grasping, and hand releasing.

21. The device of claim 1 , wherein the one or more processors to configure to synthesize the modified sound field produced by the array of loudspeakers to change a volume of the modified sound field or control a volume of a beam in the modified sound field.

22. The device of claim 1 , wherein the one or more processors are, based on the representation of the gesture, configured to synthesize the modified sound field produced by the array of loudspeakers to change a beam width of the modified sound field or change an echo depth in time of the modified sound field or change in dynamic range expansion or compression of the modified sound field.

23. The device of claim 1 , wherein the one or more processors are, based on the representation of the gesture, configured to synthesize the modified sound field produced by the array of loudspeakers to create or delete a sound null in an indicated direction relative to an axis of the array of the loudspeaker.

24. The device of claim 1 , wherein the at least one recognized movement of the user represents a sequence of two or more gestures, wherein the one or more processors are, based on the representation of the sequence of two or more gestures, configured to synthesize the modified sound field produced by the array of loudspeakers for menu navigation.

25. The device of claim 1 , wherein the one or more processors are, based on the representation of the gesture, configured to synthesize the modified sound field produced by the array of loudspeakers for user-interface feedback via sound.

26. The device of claim 1 , wherein the one or more processors are, based on the representation of the gesture configured to provide for a user-interface feedback via a display icon.

27. The device of claim 1 , wherein the representation, of the gesture that is appropriate includes for the current context is to ignore the representation of the gesture to block sound from a direction when a system is already in a blocked state in that direction.

28. The device of claim 1 , wherein the representation, of the gesture that is appropriate; includes for the current context to indicates whether the command is applied locally or globally.

29. The device of claim 1 , wherein the indicated change includes to change a beam direction of the modified sound field.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2024
From: XIANG, PEI; VISSER, ERIK
To: QUALCOMM INCORPORATED
Reel/Frame 066394/0542 →
Continuity (4)
Continuation 16586892 · Sep 27, 2019
Continuation 13775720 · Feb 25, 2013
Provisional Application 61619202 · Apr 2, 2012
Related Publication 20240098420A1 · Mar 21, 2024
References Cited (148)
US 2636943A · Schaeffer · 1953 [cited by applicant]
US 4133977A · McGuire et al. · 1979 [cited by applicant]
US 5774591A · Black et al. · 1998 [cited by applicant]
US 5796843A · Inanaga et al. · 1998 [cited by applicant]
US 6351222B1 · Swan et al. · 2002 [cited by applicant]
US 6494363B1 · Roger et al. · 2002 [cited by applicant]
US 7146011B2 · Yang et al. · 2006 [cited by applicant]
US 7184952B2 · Hillis et al. · 2007 [cited by applicant]
US 7194094B2 · Horrall et al. · 2007 [cited by applicant]
US 7277550B1 · Avendano et al. · 2007 [cited by applicant]
US 7298871B2 · Lee et al. · 2007 [cited by applicant]
US 7505898B2 · Hillis et al. · 2009 [cited by applicant]
US 7567847B2 · Basson et al. · 2009 [cited by applicant]
US 8019431B2 · Nie et al. · 2011 [cited by applicant]
US 8107639B2 · Moeller et al. · 2012 [cited by applicant]
US 8140326B2 · Chen et al. · 2012 [cited by applicant]
US 8428272B2 · Tohyama et al. · 2013 [cited by applicant]
US 9268404B2 · Clavin et al. · 2016 [cited by applicant]
US 10448161B2 · Xiang et al. · 2019 [cited by applicant]
US 20010021259A1 · Horrall · 2001 [cited by applicant]
US 20020167862A1 · Tomasi et al. · 2002 [cited by applicant]
US 20030091199A1 · Horrall et al. · 2003 [cited by applicant]
US 20030142833A1 · Roy et al. · 2003 [cited by applicant]
US 20030144848A1 · Roy et al. · 2003 [cited by applicant]
US 20040076271A1 · Koistinen et al. · 2004 [cited by applicant]
US 20040125922A1 · Specht · 2004 [cited by applicant]
US 20050065778A1 · Mastrianni et al. · 2005 [cited by applicant]
US 20050132420A1 · Howard et al. · 2005 [cited by applicant]
US 20060098830A1 · Roeder et al. · 2006 [cited by applicant]
US 20060140420A1 · Machida · 2006 [cited by applicant]
US 20060206221A1 · Metcalf · 2006 [cited by applicant]
US 20060247919A1 · Specht et al. · 2006 [cited by applicant]
US 20060247924A1 · Hillis et al. · 2006 [cited by applicant]
US 20060277039A1 · Vos et al. · 2006 [cited by applicant]
US 20070211023A1 · Boillot · 2007 [cited by applicant]
US 20070239295A1 · Thompson et al. · 2007 [cited by applicant]
US 20070263889A1 · Melanson · 2007 [cited by applicant]
US 20070269062A1 · Rodigast et al. · 2007 [cited by applicant]
US 20080101616A1 · Melchior et al. · 2008 [cited by applicant]
US 20080126086A1 · Vos et al. · 2008 [cited by applicant]
US 20080130923A1 · Freeman · 2008 [cited by applicant]
US 20080235008A1 · Ito et al. · 2008 [cited by applicant]
US 20090060236A1 · Johnston et al. · 2009 [cited by applicant]
US 20090074199A1 · Kierstein et al. · 2009 [cited by applicant]
US 20090102800A1 · Keenan · 2009 [cited by applicant]
US 20090195518A1 · Mattice et al. · 2009 [cited by applicant]
US 20090304205A1 · Hardacker et al. · 2009 [cited by applicant]
US 20100098275A1 · Metcalf · 2010 [cited by applicant]
US 20100158263A1 · Katzer et al. · 2010 [cited by applicant]
US 20100182231A1 · Morimiya et al. · 2010 [cited by applicant]
US 20100202656A1 · Ramakrishnan et al. · 2010 [cited by applicant]
US 20100208912A1 · Tohyama et al. · 2010 [cited by applicant]
US 20100226499A1 · De Bruijn et al. · 2010 [cited by applicant]
US 20100241999A1 · Russ et al. · 2010 [cited by applicant]
US 20110038489A1 · Visser et al. · 2011 [cited by applicant]
US 20110063442A1 · Aarts et al. · 2011 [cited by applicant]
US 20110096941A1 · Marzetta et al. · 2011 [cited by applicant]
US 20110103620A1 · Strauss et al. · 2011 [cited by applicant]
US 20110182438A1 · Koike et al. · 2011 [cited by applicant]
US 20110242305A1 · Peterson et al. · 2011 [cited by applicant]
US 20110254762A1 · Dahl et al. · 2011 [cited by applicant]
US 20110289455A1 · Reville et al. · 2011 [cited by applicant]
US 20120005632A1 · Broyles, III et al. · 2012 [cited by applicant]
US 20120014525A1 · Ko et al. · 2012 [cited by applicant]
US 20120020480A1 · Visser et al. · 2012 [cited by applicant]
US 20120053931A1 · Holzrichter · 2012 [cited by applicant]
US 20120114137A1 · Tsurumi · 2012 [cited by applicant]
US 20120120073A1 · Haker et al. · 2012 [cited by applicant]
US 20120194561A1 · Grossinger et al. · 2012 [cited by applicant]
US 20120265534A1 · Coorman et al. · 2012 [cited by applicant]
US 20130106686A1 · Bennett · 2013 [cited by applicant]
US 20130121515A1 · Hooley et al. · 2013 [cited by applicant]
US 20130223658A1 · Betlehem et al. · 2013 [cited by applicant]
US 20130259238A1 · Xiang et al. · 2013 [cited by applicant]
US 20130259254A1 · Xiang et al. · 2013 [cited by applicant]
US 20130315413A1 · Yamakawa et al. · 2013 [cited by applicant]
US 20140006017A1 · Sen · 2014 [cited by applicant]
US 20140086426A1 · Yamakawa et al. · 2014 [cited by applicant]
US 20140328487A1 · Hiroe · 2014 [cited by applicant]
US 20140337016A1 · Herbig et al. · 2014 [cited by applicant]
US 20200077193A1 · Xiang et al. · 2020 [cited by applicant]
CN 1347263A · 2002 [cited by applicant]
CN 101313518A · 2008 [cited by applicant]
CN 101794180A · 2010 [cited by applicant]
CN 102027440A · 2011 [cited by applicant]
CN 102117117A · 2011 [cited by applicant]
JP H05241573A · 1993 [cited by applicant]
JP 2008103851A · 2008 [cited by applicant]
JP 2010045432A · 2010 [cited by applicant]
KR 20120006710A · 2012 [cited by applicant]
WO 9948085A1 · 1999 [cited by applicant]
WO 2009156928A1 · 2009 [cited by applicant]
WO 2011036618A2 · 2011 [cited by applicant]
WO 2011059202A2 · 2011 [cited by applicant]
WO 2011135283A2 · 2011 [cited by applicant]
WO 2012015843A1 · 2012 [cited by applicant]
Yoo, S.: “Speech Decomposition and Enhancement,” University of Pittsburgh, 2005, p. 178, Section 1.1, p. 1-3, Section 2.1.2 p. 6, Section 2.2.3 p. 10-11, Fig.2, Section 3.2, Section 5.1.1, p. 88, Para.2, Section 6.1 p. … [cited by applicant]
Zhang Z., “Human Body Language Understanding with 3D Sensors,” 2011, 54 pages. [cited by applicant]
Acero A., “Audio and Video Research in Kinect,” Microsoft Research, Jun. 2011, 58 pages. [cited by applicant]
Atlas L., et al. “Applications and Justification of Coherent Modulation Filtering”, pp. 7, Accessed online Apr. 2, 2013 at www.silicon-speech.com/Media/TemporalDynamics/PDF/Atlas_TemporalDynamics.pdf. [cited by applicant]
Atlas L., et al., “Joint Acoustic and Modulation Frequency”, EURASIP J. Appl. Sig, Proc, 2003, vol. 7, pp. 668-675. [cited by applicant]
Belgraver Thissen W.P.C., “A Comparative Study of Optical Depth Sensors for User Interaction”, Master project report, Eindhoven University of Technology, Jul. 2011, 93 pages. [cited by applicant]
Bergh M.V.D., et al., “Real-time 3D Hand Gesture Interaction with a Robot for Understanding Directions from Humans,” IEEE, 2011, pp. 357-362. [cited by applicant]
Caputo M., et al., “3D Hand Gesture Recognition Based on Sensor Fusion of Commodity Hardware,” in Proceedings of Mensch Computer, 2012, pp. 293-302. [cited by applicant]
Chang J.S., et al., “Vision-Based Interface for Integrated Home Entertainment System”, Springer-Verlag Berlin Heidelberg, pp. 176-183, (Year: 2005). [cited by applicant]
Do C.T., et al., “On the Recognition of Cochlear Implant-Like Spectrally Reduced Speech With MFCC and HMM-Based ASR”, published by IEEE Transactions on Audio, Speech and Language Processing on Jul. 1, 2010 at IEEE Servi… [cited by applicant]
Doliotis P., et al., “Comparing Gesture Recognition Accuracy Using Color and Depth Information,” PETRA '11 Proceedings of the 4th International Conference on Pervasive Technologies Related to Assistive Environments, 201… [cited by applicant]
Dondi P., et al., “Gesture Recognition by Data Fusion of Time-of-Flight and Color Cameras”, World Academy of Science, Engineering and Technology 59, 2011, pp. 1954-1959. [cited by applicant]
Drake A., “Kinect Hand Recognition and Tracking,” Department of Computer Science Engineering, 2012, 5 pages. [cited by applicant]
Du H., et al., “Hand Gesture Recognition Using Kinect,” Dec. 15, 2011, Technical Report No. ECE-2011-04, pp. 1-23. [cited by applicant]
Elgendi M., et al., “Real-Time Speed Detection of Hand Gesture using Kinect”, Workshop on Autonomous Social Robots and Virtual Humans, the 25th Annual Conference on Computer Animation and Social Agents (CASA 2012), Sing… [cited by applicant]
Frati V., et al., “Using Kinect for hand tracking and rendering in wearable haptics,” IEEE World Haptics Conference, 2011, pp. 317-321. [cited by applicant]
Grossinger., et al., “U.S. Appl. No. 61/244,473”. [cited by applicant]
Gunes H., et al., “Automatic Visual Recognition of Face and Body Action Units”, Third International Conference on Information Technology and Applications (ICITA'05), DOI:10.1109/icita.2005.83, (Year: 2005) 6 Pages. [cited by applicant]
Guo J., “Hand Gesture Recognition and Interaction with 3D stereo Camera,” Nov. 2011, 34 pages. [cited by applicant]
Hall J.C., “How to do Gesture Recognition with Kinect Using Hidden Markov Models (HMMs),” Dec. 22, 2011, Creative Distraction, 12 pages, [Retrieved on Oct. 24, 2012]. [cited by applicant]
Hamlynkinect, “Hand Detection Algorithms,” Retrieved on Oct. 31, 2012, 2 pages, URL: http://hamlynkinect.wikispaces.com/Hand+Detection+Algorithms. [cited by applicant]
International Preliminary Report on Patentability—PCT/US2013/029038—The International Bureau of WIPO—Geneva, Switzerland, Jul. 18, 2014. [cited by applicant]
International Preliminary Report on Patentability—PCT/US2013/033082, The International Bureau of WIPO—Geneva, Switzerland, Jul. 11, 2014. [cited by applicant]
International Preliminary Report on Patentability—PCT/US2013/043341, The International Bureau of WIPO—Geneva, Switzerland, Oct. 24, 2014. [cited by applicant]
International Search Report and Written Opinion—PCT/US2013/029038—ISA/EPO—Jun. 4, 2013. [cited by applicant]
International Search Report and Written Opinion—PCT/US2013/033082—ISA/EPO—Jun. 25, 2013. [cited by applicant]
International Search Report and Written Opinion—PCT/US2013/043341—ISA/EPO—Feb. 27, 2014. [cited by applicant]
Kawahara H., et al., “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based FO extraction: Possible role of a repetitive structure in sounds”, Speech C… [cited by applicant]
Kurakin A., et al., “A Real Time System For Dynamic Hand Gesture Recognition With A Depth Sensor,” 20th European Signal Processing Conference (EUSIPCO) 2012, pp. 1975-1979. [cited by applicant]
Langton C., “Signal Processing Simulation Newsletter”, 1999, 11 Pages, http://complextoreal.com/wpcontent/uploads/2013/01/tcomplex.pdf as of Apr. 1, 2015. [cited by applicant]
Li X., et al., “Harmonic Coherent Demodulation for Improving Sound Coding in Cochlear Implants”, Acoustics Speech and Signal Processing (ICASSP), 2010, IEEE International Conference on, Mar. 14-19, 2010, pp. 5462-5465. [cited by applicant]
Li Y., “Hand Gesture Recognition Using Kinect,” Aug. 2012, 44 pages. [cited by applicant]
Liu N., et al., “Gesture Classification Using Hidden Markov Models and Viterbi Path Counting,” Proc, VIIth Digital Image Computing: Techniques and Applications, Dec. 10-12, 2003, pp. 273-282. [cited by applicant]
Mahmoodzadeh A., et al., “Single channel speech separation in modulation frequency domain based on a novel pitch range estimation method”, EURASIP J. Advances in Sig. Proc, 2012, vol. 67, pp. 10. [cited by applicant]
Mizoguchi, et al., “Invisible Messenger: Visually Steerable Sound Beam Forming System based on Face Tracking and Speaker Array,” SICE Annual Conference in Fukui, Fukui University, Japan, Aug. 4-6, 2003, pp. 3007-3011. [cited by applicant]
Park S., et al., “3D hand tracking using Kalman filter in depth space,” EURASIP Journal on Advances in Signal Processing, 2012, vol. 36, pp. 1-18, URL: http://asp.eurasipjournals.com/content/2012/1/36. [cited by applicant]
Pearse S., “Gestural Mappings: Towards the Creation of a Three Dimensional Composition Environment”, In Proceedings of the International Computer Music Conference, pp. 126-129, 2011. [cited by applicant]
Rafaely, et al., “Optimal Model-Based Beamforming and Independent Steering for Spherical Loudspeaker Arrays,” IEEE Transactions on audio, speech, and language processing, vol. 19. 19, No. 7, Sep. 2011, pp. 2234-2238. [cited by applicant]
Ren Z., et al., “Depth Camera Based Hand Gesture Recognition and its Applications in Human-Computer-Interaction,” 8th International Conference on Information, Communications and Signal Processing (ICICS), 2011, pp. 1-5. [cited by applicant]
Ren Z., et al., “Robust Hand Gesture Recognition with Kinect Sensor,” Proceedings of the 19th ACM international conference on Multimedia, 2011, pp. 759-760. [cited by applicant]
Schimmel S.M. et al., “Coherent Envelope Detector for Modulation Filtering of Speech”, in Proceedings of ICASSP, vol. 1, pp. 221-224, Philadelphia, USA, May 2005. [cited by applicant]
Schimmel S.M., et al., “Feasibility of Single Channel Speaker Separation Based on Modulation Frequency Analysis”, Acoustics, Speech and Signal Processing, 2007. ICASSP 2007. IEEE International Conference on, Apr. 15-20,… [cited by applicant]
Schimmel S.M., et al., “Frequency Reassignment for Coherent Modulation Filtering”, Proceedings of ICASSP'06, 2006, pp. 261-264. [cited by applicant]
Schmeder, “An Exploration of Design Parameters for Human Interactive Systems with Compact Sphere Loudspeaker Arrays,” Ambisonics Symposium, Jun. 2009, pp. 1-11. [cited by applicant]
Shi, et al., “Development of a Parametric Loudspeaker: A Novel Directional Sound Generation Technology,” IEEE Potentials, Nov.-Dec. 2010, vol. 29 Issue 6, pp. 20-24. [cited by applicant]
Shimoda H., et al., “A Study on Real-time Gesture Classification Method,” 2001, 5 pages. [cited by applicant]
Steinberg I., et al., “Hand Gesture Recognition in Images and Video,” Irwin and Joan Jacobs Center for Communication and Information Technologies, CCIT Report #763, Mar. 2010, pp. 1-20. [cited by applicant]
Tang M., “Hand Gesture Recognition Using Microsoft's Kinect,” Mar. 16, 2011, pp. 1-7. [cited by applicant]
Tang M., “Recognizing Hand Gestures with Microsoft's Kinect,” 2011, pp. 1-5. [cited by applicant]
Trigo T.R., et al., “An Analysis of Features for Hand-Gesture Classification,” IWSSIP 2010—17th International Conference on Systems, Signals and Image Processing, 2010, pp. 412-415. [cited by applicant]
Trindade P., et al., “Hand gesture recognition using color and depth images enhanced with hand angular pose data,” 2012 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI), … [cited by applicant]
Xu W., et al., “Gesture Recognition based on 2D and 3D Feature by using Kinect Device,” 2012, pp. 77-79. [cited by applicant]