IP Library Granted Patent US 11,272,125
Granted Patent B2
US 11,272,125 · App. 17/129,258 · Granted Mar 8, 2022

Systems and methods for automatic detection and insetting of digital streams into a video

Inventors: Yulius Tjahjadi (San Mateo, CA); Donald Kimber (Foster City, CA); Qiong Liu (Cupertino, CA); Laurent Denoue (Verona, IT)
Assignee: FUJIFILM Business Innovation Corp.
H04N5/272H04N5/23238
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,272,125
App. No.
17/129,258
Granted
Mar 8, 2022
Kind
B2
Abstract

Systems and methods for automatic detection and insetting of digital streams into a video. Various examples of such regions of interest to the user include, without limitation, content displayed on various electronic displays or written on electronic paper (electronic ink), content projected on various surfaces using electronic projectors, content of paper documents appearing in the video and/or content written on white (black) boards inside the video. For some content, such as whiteboards, paper documents or paintings in a museum, a participant (or curator) could have taken pictures of the regions, again stored digitally somewhere and available for download. These digital streams with content of interest to the user are obtained and then inset onto the view generated from the raw video feed, giving users the ability to view them at their native high resolution.

Claims (37)

1. A system comprising:

at least one camera for acquiring a video of an environment which includes objects comprising surfaces, including at least one surface where a first digital media content is displayed; and

at least one processor executing instruction stored in a memory to:

detect the surfaces from the video using object recognition;

identify one or more surfaces, of the detected surfaces, within the video as one or more inset candidate regions; and

inset a second digital media content into the identified one or more inset candidate regions, the second digital media content representative of the first digital media content.

2. The system of claim 1 , wherein the one or more surfaces identified as one or more inset candidate regions within the video is a display screen.

3. The system of claim 1 , wherein the one or more surfaces identified as one or more inset candidate regions within the video is a whiteboard.

4. The system of claim 1 , wherein the first and second digital media content is an image.

5. The system of claim 1 , wherein the first and second digital media content is a video stream.

6. The system of claim 1 , wherein a resolution of the second digital media content is higher than the resolution of the video.

7. The system of claim 1 , wherein the at least one processor executes the instruction to:

detect an occlusion of the identified one or more inset candidate regions by an object between the one or more surfaces identified as the insert candidate regions and the at least one camera; and

cut the inset second digital media content based on the detected occlusion.

8. The system of claim 7 , wherein the at least one processor executes the instruction to:

compute a mask based on the detected occlusion; and

cut the inset second digital media content using the mask.

9. The system of claim 1 , wherein the one or more inset candidate regions are identified based on a position of the at least one camera in relation to a position of the inset candidate region.

10. The system of claim 1 , wherein the one or more inset candidate regions are additionally identified based on an input.

11. The system of claim 1 , wherein the second digital media content to be inset into the identified one or more inset candidate regions is selected based on location of the one or more surfaces identified as the one or more inset candidate regions within the video.

12. The system of claim 1 , wherein the at least one camera and the identified one or more surfaces are fixed.

13. The system of claim 1 , wherein the one or more inset candidate regions are manually adjustable.

14. The system of claim 1 , wherein the at least one processor executes the instruction to identify a plurality of surfaces, of the detected surfaces, within the video as a plurality of inset candidate regions and select at least one digital media content to be inset into each of the plurality of inset candidate regions.

15. A method comprising:

acquiring a video of an environment includes objects comprising surfaces, including at least one surface where a first digital media content is displayed, using at least one camera;

detecting the surfaces from the video using object recognition;

identifying one or more surfaces, of the detected surfaces, within the video as one or more inset candidate regions; and

insetting a second digital media content into the identified one or more inset candidate regions, the second digital media content representative of the first digital media content.

16. The method of claim 15 , wherein the at least one camera and the identified one or more surfaces are fixed.

17. The method of claim 15 , wherein the one or more inset candidate regions are manually adjustable.

18. A non-transitory computer-readable medium embodying a set of instructions implementing a method comprising:

acquiring a video of an environment includes objects comprising surfaces, including at least one surface where a first digital media content is displayed, using at least one camera;

detecting the surfaces from the video using object recognition;

identifying one or more surfaces, of the detected surfaces, within the video as one or more inset candidate regions; and

insetting a second digital media content into the identified one or more inset candidate regions, the second digital media content representative of the first digital media content.

19. The non-transitory computer-readable medium of claim 18 , wherein the at least one camera and the identified one or more surfaces are fixed.

20. The non-transitory computer-readable medium of claim 18 , wherein the one or more inset candidate regions are manually adjustable.

Assignments (1)
CHANGE OF NAME Recorded May 25, 2021
From: FUJI XEROX CO., LTD.
To: FUJIFILM BUSINESS INNOVATION CORP.
Reel/Frame 056392/0541 →
Continuity (2)
Continuation 16031068 · Jul 10, 2018
Related Publication 20210112209A1 · Apr 15, 2021