IP Library Granted Patent US 12,335,060
Granted Patent B1
US 12,335,060 · App. 18/143,441 · Granted Jun 17, 2025

Audio focus in a virtual meeting based on eye tracking

Inventors: Robert Allen Ryskamp (Mountain View, CA); Adam Justin Spooner (Greensboro, NC)
Assignee: Zoom Communications, Inc.
H04L12/1831G06F3/013
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,335,060
App. No.
18/143,441
Granted
Jun 17, 2025
Kind
B1
Abstract

One example method includes joining, from a client device, a virtual meeting hosted by a virtual meeting provider, the virtual meeting comprising a plurality of participants, displaying the virtual meeting on the display of the client device, receiving an eye tracking signal from an eye tracking sensor, the eye tracking signal associated with a first user of the client device, and determining, based at least in part on the eye tracking signal, a location of the display on which the first user is focused. The method further includes identifying a first audio stream of a plurality of audio streams associated with the virtual meeting based at least in part on the location, and responsive to identifying the first audio stream, increasing a volume of the first audio stream relative to the rest of the plurality of audio streams.

Claims (43)

1. A method comprising:

joining, from a client device, a virtual meeting hosted by a virtual meeting provider, the virtual meeting comprising a plurality of participants;

displaying the virtual meeting on the display of the client device;

receiving an eye tracking signal from an eye tracking sensor, the eye tracking signal associated with a first user of the client device;

autocorrecting the eye tracking sensor, where autocorrecting further comprises triangulating between the position of the eye tracking sensor, the location of the display on which the first user is focused, and a position of the first user;

determining, based at least in part on the eye tracking signal, a location of the display on which the first user is focused;

identifying a first audio stream of a plurality of audio streams associated with the virtual meeting based at least in part on the location; and

responsive to identifying the first audio stream, increasing a volume of the first audio stream relative to the rest of the plurality of audio streams.

2. The method of claim 1 , further comprising identifying a second audio stream of the plurality of audio streams, the second audio stream associated with the first audio stream.

3. The method of claim 2 , wherein the first audio stream and the second audio stream are associated with a first user group of a plurality of user groups in the virtual meeting.

4. The method of claim 1 , further comprising calibrating the eye tracking sensor.

5. The method of claim 4 , wherein calibrating the eye tracking sensor further comprises displaying a grid on the display, the grid comprising a plurality of visual identifiers each associated with one cell in the grid.

6. The method of claim 1 , further comprising:

determining that a second user associated with the first audio stream is focusing on the first user, and

responsive to the determination, increasing the volume of the first audio stream based on the determination.

7. A system comprising:

a non-transitory computer-readable medium;

a communications interface; and

one or more processors communicatively coupled to the non-transitory computer-readable medium and the communications interface, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

join, from a client device, a virtual meeting hosted by a virtual meeting provider, the virtual meeting comprising a plurality of participants;

display the virtual meeting on the display of the client device;

receive an eye tracking signal from an eye tracking sensor, the eye tracking signal associated with a first user of the client device;

autocorrect the eye tracking sensor, wherein the processor-executable instructions to autocorrect further comprise processor-executable instructions to triangulate between the position of the eye tracking sensor, the location of the display on which the first user is focused, and a position of the first user;

determine, based at least in part on the eye tracking signal, a location of the display on which the first user is focused;

identify a first audio stream of a plurality of audio streams associated with the virtual meeting based at least in part on the location; and

responsive to identifying the first audio stream, increase a volume of the first audio stream relative to the rest of the plurality of audio streams.

8. The system of claim 7 , further comprising processor-executable instructions stored in the non-transitory computer-readable medium to identify a second audio stream of the plurality of audio streams, the second audio stream associated with the first audio stream.

9. The system of claim 8 , wherein the first audio stream and the second audio stream are associated with a first user group of a plurality of user groups in the virtual meeting.

10. The system of claim 7 , further comprising processor-executable instructions stored in the non-transitory computer-readable medium to calibrate the eye tracking sensor.

11. The system of claim 10 , wherein calibrating the eye tracking sensor further comprises displaying a grid on the display, the grid comprising a plurality of visual identifiers each associated with one cell in the grid.

12. The system of claim 7 , further comprising processor-executable instructions stored in the non-transitory computer-readable medium to:

determine that a second user associated with the first audio stream is focusing on the first user, and

responsive to the determination, increase the volume of the first audio stream based on the determination.

13. A non-transitory computer-readable medium comprising processor-executable instructions configured to cause a processor to:

join, from a client device, a virtual meeting hosted by a virtual meeting provider, the virtual meeting comprising a plurality of participants;

display the virtual meeting on the display of the client device;

receive an eye tracking signal from an eye tracking sensor, the eye tracking signal associated with a first user of the client device;

autocorrect the eye tracking sensor, wherein the processor-executable instructions to autocorrect further comprise processor-executable instructions to triangulate between the position of the eye tracking sensor, the location of the display on which the first user is focused, and a position of the first user;

determine, based at least in part on the eye tracking signal, a location of the display on which the first user is focused;

identify a first audio stream of a plurality of audio streams associated with the virtual meeting based at least in part on the location; and

responsive to identifying the first audio stream, increase a volume of the first audio stream relative to the rest of the plurality of audio streams.

14. The non-transitory computer-readable medium of claim 13 , further comprising processor-executable instructions configured to cause a processor to identify a second audio stream of the plurality of audio streams, the second audio stream associated with the first audio stream.

15. The non-transitory computer-readable medium of claim 14 , wherein the first audio stream and the second audio stream are associated with a first user group of a plurality of user groups in the virtual meeting.

Assignments (1)
CHANGE OF NAME Recorded May 19, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 071308/0893 →
Continuity (1)
Continuation 17876701 · Jul 29, 2022
References Cited (7)
US 10581625B1 · Pandey · 2020 [cited by examiner]
US 20070299710A1 · Haveliwala · 2007 [cited by examiner]
US 20140313124A1 · Kim · 2014 [cited by examiner]
US 20160343164A1 · Urbach · 2016 [cited by examiner]
US 20230231983A1 · Troje · 2023 [cited by examiner]
US 20230298281A1 · Pease · 2023 [cited by examiner]
US 20240031531A1 · Krol · 2024 [cited by examiner]
Cited By (2)
US 12,688,005 US 12,726,373