IP Library › Granted Patent US 11,363,078
Granted Patent B2
US 11,363,078 · App. 17/328,532 · Granted Jun 14, 2022

System and method for augmented reality video conferencing

Inventors: Gaurang Bhatt (Herndon, VA); Negar Kalbasi (Washington, DC)
Assignee: CAPITAL ONE SERVICES, LLC
H04L65/4015G06T11/00G06V20/40G06V40/172G10L17/06H04L67/38H04N7/155
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,363,078
App. No.
17/328,532
Granted
Jun 14, 2022
Kind
B2
Abstract

A system includes a plurality of capturing devices and a plurality of displaying devices. The capturing devices and the displaying devices can be communicatively connected to a server. The server can receive captured data from the capturing devices and transform the data into a digitized format. At least one of the capturing devices can record a video during a meeting and the capturing device can transmit the captured video to the server as a video feed. The video feed can show an area that includes handwritten text, e.g., a whiteboard. The server can receive the video feed from the capturing device and perform various processes on the video. For example, the server can perform a voice recognition, text recognition, handwriting recognition, face recognition and/or object recognition technique on the captured video.

Claims (43)

1. A method comprising:

receiving, at a server, a first video feed from a first camera;

processing, using a processor of the server, the first video feed to:

extract a first text from a first segment of the first video feed; and

identify a first person;

receiving, at the server, a second video feed from a second camera;

processing, using the processor of the server, the second video feed to:

extract a second text from a second segment of the second video feed; and

identify a second person;

generating, using the processor of the server, a slide showing a first summary of the first text and a second summary of the second text, wherein:

the first summary is displayed in association with the first person and the second summary is displayed in association with the second person; and

the first summary is selectable to display a first media file and the second summary is selectable to display a second media file.

2. The method of claim 1 , further comprising processing, using the processor of the server, the first video feed to extract a first shape from the first segment of the first video feed.

3. The method of claim 2 , wherein the processor extracts the first shape using a shape recognition module.

4. The method of claim 2 , wherein the slide is generated to show the first shape in digital format.

5. The method of claim 1 , wherein the slide further displays the first text or the second text.

6. The method of claim 5 , wherein the processor generates the first summary using a machine learning technique.

7. The method of claim 5 , wherein the processor generates the first summary using a natural language processing technique.

8. The method of claim 1 , wherein the processor identifies the first person using a face recognition technique.

9. The method of claim 1 , wherein the processor identifies the first person using a voice recognition technique.

10. The method of claim 1 , wherein the first media file is the first segment of the first video feed.

11. The method of claim 1 , wherein the first media file is a voice recording.

12. The method of claim 1 , wherein the second media file is the second segment of the second video feed.

13. The method of claim 1 , wherein the first summary is selectable to display an identity of the first person.

14. The method of claim 1 , wherein the slide is included in a PDF document or a Word document.

15. The method of claim 1 , wherein the slide is text-searchable.

16. The method of claim 1 , wherein the processor extracts the first text from the first segment of the first video feed using a transcription module.

17. The method of claim 1 , wherein the processor extracts the first text from the first segment of the first video feed using a text recognition module.

18. The method of claim 1 , wherein the processor extracts the first text from the first segment of the first video feed using a handwriting recognition module.

19. The method of claim 1 , wherein the processor identifies the first person using a facial recognition module configured to:

scan the first video feed;

detect a face in the first video feed; and

identify the face based on a comparison of the face with a plurality of photos of faces stored in a database.

20. A method comprising:

receiving, at a server, a first video feed from a first camera;

processing, using a processor of the server, the first video feed to:

subtract a background from the first video feed to generate an overlay feed; and

identify a first person;

receiving, at the server, a second video feed from a second camera of an AR glass;

processing, using the processor of the server, the second video feed to:

extract a second text from a second segment of the second video feed; and

identify a second person; and

transmit, using the processor of the server, the overlay feed to the AR glass, wherein the overlay feed is configured to be superimposed over a field of view of the AR glass.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2021
From: BHATT, GAURANG; KALBASI, NEGAR
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 056331/0884 →
Continuity (2)
Continuation 16991286 · Aug 12, 2020
Related Publication 20220053039A1 · Feb 17, 2022