IP Library Granted Patent US 12,651,455
Granted Patent B2
US 12,651,455 · App. 17/482,402 · Granted Jun 9, 2026

Capturing objects in an unstructured video stream

Inventor: Ian M. Richter (Los Angeles, CA)
Assignee: APPLE INC.
G06V20/20G06F16/75G06F16/7837G06F16/7867G06T7/11G06V10/17G06V10/25G06V10/82G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,455
App. No.
17/482,402
Granted
Jun 9, 2026
Kind
B2
Abstract

A method includes obtaining a first unstructured video stream that provides pixel values for a plurality of pixels and corresponds to a portion of a second unstructured video stream being displayed on a second electronic device different from the first electronic device. Obtaining the first unstructured video stream includes obtaining pass-through image data including the portion of a second unstructured video stream. The method includes generating respective pixel characterization vectors for a first portion of the plurality of pixels. Generating each of the respective pixel characterization vectors includes determining a respective instance label value. The method includes identifying a first object within the first portion of the plurality of pixels associated with a particular instance label value. The method includes generating respective semantic label values corresponding to pixels associated with the first object. The respective semantic label values are added to pixel characterization vectors associated with the first object.

Claims (48)

1 . A method comprising:

at a first electronic device including one or more processors, one or more image sensors, and a non-transitory memory:

obtaining, using the one or more image sensors, a first unstructured video stream that provides pixel values for a plurality of pixels for rendering in an extended reality (XR) environment, wherein the one or more image sensors capture pass-through image data including a portion of a second unstructured video stream being displayed on a secondary display of a second electronic device that is different from the first electronic device;

generating respective pixel characterization vectors for a portion of the plurality of pixels, wherein each of the respective pixel characterization vectors is associated with a corresponding pixel in the portion of the plurality of pixels, wherein generating each of the respective pixel characterization vectors includes determining a respective instance label value and adding the respective instance label value to each of the respective pixel characterization vectors indicating separate objects in one or more images of the first unstructured video stream;

identifying a first object within the portion of the plurality of pixels based on satisfying an object confidence threshold that indicates that pixels are associated with pixel characterization vectors that each include a first instance label value that is associated with an object represented in the portion of the plurality of pixels;

generating respective semantic label values corresponding to pixels associated with the first object, wherein the respective semantic label values are added to pixel characterization vectors associated with the first object and characterize the first object; and

providing a first XR affordance corresponding to the identified first object to instantiate an objective-effectuator, wherein the objective-effectuator performing actions in the XR environment is characterized by the respective semantic label values.

2 . The method of claim 1 , further comprising appending the respective semantic label values to the pixel characterization vectors associated with the first object.

3 . The method of claim 1 , further comprising displaying, via a primary display of the first electronic device, extended reality (XR) content that corresponds to the first object in the first unstructured video stream, wherein the XR content is displayed overlaid on the pass-through image data that include the portion of the second unstructured video stream.

4 . The method of claim 1 , wherein the respective pixel characterization vectors are generated by an instance segmentation classifier.

5 . The method of claim 1 , further comprising:

identifying a second object within the portion of the plurality of pixels associated with a second instance label value that is different from the first instance label value; and

generating additional semantic label values corresponding to pixels associated with the second object in the first unstructured video stream, wherein the additional semantic label values are added to the pixel characterization vectors associated with the second object.

6 . The method of claim 1 , wherein the first electronic device and the second electronic device are separate from each other.

7 . The method of claim 3 , wherein the XR content is based on the respective semantic label values corresponding to the pixels associated with the first object.

8 . The method of claim 3 , further comprising, in response to obtaining, from one or more input devices, a first input corresponding to the first XR affordance:

in accordance with a determination that the first input corresponds to a first input type, displaying, via the primary display, informational XR content corresponding to the first object, wherein the informational XR content is based on the respective semantic label values corresponding to the pixels associated with the first object; and

in accordance with a determination that the first input corresponds to a second input type different from the first input type, displaying, via the primary display, the objective- effectuator based on the respective semantic label values corresponding to the pixels associated with the first object, wherein the objective-effectuator is characterized by a set of predefined objectives and a set of visual rendering attributes.

9 . The method of claim 4 , wherein the respective semantic label values are generated by a semantic segmentation classifier that is different from the instance segmentation classifier.

10 . The method of claim 8 , wherein displaying the objective-effectuator includes instantiating, the objective-effectuator in an emergent content container characterized by contextual information, wherein the emergent content container enables the objective-effectuator to perform a set of actions that satisfy the set of predefined objectives.

11 . The method of claim 10 , further comprising:

displaying, via the primary display, a second XR affordance in association with the emergent content container, wherein the second XR affordance controls an operation of the emergent content container; and

in response to detecting, via the one or more input devices, a second input corresponding to the second XR affordance, modifying the objective-effectuator.

12 . The method of claim 10 , further comprising:

generating a sequence of actions of the set of actions based on the contextual information and a particular objective of the set of predefined objectives; and

modifying, via the primary display, the objective-effectuator based on the sequence of actions.

13 . A first electronic device comprising:

one or more processors;

one or more image sensors;

a non-transitory memory; and

one or more programs, wherein the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

obtaining, using the one or more image sensors, a first unstructured video stream that provides pixel values for a plurality of pixels for rendering in an extended reality (XR) environment, wherein the one or more image sensors capture pass-through image data including a portion of a second unstructured video stream being displayed on a secondary display of a second electronic device that is different from the first electronic device;

generating respective pixel characterization vectors for a portion of the plurality of pixels, wherein each of the respective pixel characterization vectors is associated with a corresponding pixel in the portion of the plurality of pixels, wherein generating each of the respective pixel characterization vectors includes determining a respective instance label value and adding the respective instance label value to each of the respective pixel characterization vectors indicating separate objects in one or more images of the first unstructured video stream;

identifying a first object within the portion of the plurality of pixels based on satisfying an object confidence threshold that indicates that pixels are associated with pixel characterization vectors that each include a first instance label value that is associated with an object represented in the portion of the plurality of pixels;

generating respective semantic label values corresponding to pixels associated with the first object, wherein the respective semantic label values are added to pixel characterization vectors associated with the first object and characterize the first object; and

providing a first XR affordance corresponding to the identified first object to instantiate an objective-effectuator, wherein the objective-effectuator performing actions in the XR environment is characterized by the respective semantic label values.

14 . The first electronic device of claim 13 , wherein the first electronic device includes a primary display, and wherein the one or more programs include instructions for displaying, via the primary display, extended reality (XR) content that corresponds to the first object in the first unstructured video stream, wherein the XR content is displayed overlaid on the pass-through image data that include the portion of the second unstructured video stream.

15 . The first electronic device of claim 14 , wherein the XR content is based on the respective semantic label values corresponding to the pixels associated with the first object.

16 . The first electronic device of claim 14 , wherein the one or more programs include instructions for, in response to obtaining, from one or more input devices, a first input corresponding to the first XR affordance:

in accordance with a determination that the first input corresponds to a first input type, displaying, via the primary display, informational XR content corresponding to the first object, wherein the informational XR content is based on the respective semantic label values corresponding to the pixels associated with the first object; and

in accordance with a determination that the first input corresponds to a second input type different from the first input type, displaying, via the primary display, the objective- effectuator based on the respective semantic label values corresponding to the pixels associated with the first object, wherein the objective-effectuator is characterized by a set of predefined objectives and a set of visual rendering attributes.

17 . The first electronic device of claim 16 , wherein displaying the objective- effectuator includes instantiating, the objective-effectuator in an emergent content container characterized by contextual information, wherein the emergent content container enables the objective-effectuator to perform a set of actions that satisfy the set of predefined objectives.

18 . A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which, when executed by a first electronic device including one or more processors and one or more image sensors, cause the first electronic device to:

obtain, using the one or more image sensors, a first unstructured video stream that provides pixel values for a plurality of pixels for rendering in an extended reality (XR) environment, wherein the one or more image sensors capture pass-through image data including a portion of a second unstructured video stream being displayed on a secondary display of a second electronic device that is different from the first electronic device;

generate respective pixel characterization vectors for a portion of the plurality of pixels, wherein each of the respective pixel characterization vectors is associated with a corresponding pixel in the portion of the plurality of pixels, wherein generating each of the respective pixel characterization vectors includes determining a respective instance label value and adding the respective instance label value to each of the respective pixel characterization vectors indicating separate objects in one or more images of the first unstructured video stream;

identifying a first object within the portion of the plurality of pixels based on satisfying an object confidence threshold that indicates that pixels are associated with pixel characterization vectors that each include a first instance label value that is associated with an object represented in the portion of the plurality of pixels;

generate respective semantic label values corresponding to pixels associated with the first object, wherein the respective semantic label values are added to pixel characterization vectors associated with the first object and characterize the first object; and

provide a first XR affordance corresponding to the identified first object to instantiate an objective-effectuator, wherein the objective-effectuator performing actions in the XR environment is characterized by the respective semantic label values.

Continuity (3)
Continuation PCTUS2020030025 · Apr 27, 2020
Provisional Application 62840263 · Apr 29, 2019
Related Publication 20220012283A1 · Jan 13, 2022
References Cited (9)
US 10042048B1 · Moya · 2018 [cited by examiner]
US 20140101691A1 · Sinha · 2014 [cited by examiner]
US 20160140724A1 · Ji · 2016 [cited by examiner]
US 20200051337A1 · Reynolds · 2020 [cited by examiner]
CN 102509104A · 2012 [cited by examiner]
PCT International Search Report and Written Opinion dated Aug. 11, 2020, International Application No. PCT/US2020/030025, pp. 1-12. [cited by applicant]
Simon Bergweiler et al., “Foundations of Semantic Television Design of a Distributed and Gesture-Based Television System,” International Journal on Advances in Intelligent Systems, vol. 8, No. 1 & 2, 2015, pp. 194-208. [cited by applicant]
Timothy Neate et al., “Cross-device media: a review of second screening and multi-device television,” Personal and Ubiquitous Computing, vol. 21, No. 2, 2017, pp. 391-405. [cited by applicant]
Sergio Goldenberg, “Creating augmented and immersive television experiences using a semantic framework,” Proceedings of the 1st international conference on Designing interactive user experiences for TV and video, 2008, … [cited by applicant]