IP Library › Granted Patent US 12,282,696
Granted Patent B2
US 12,282,696 · App. 18/723,429 · Granted Apr 22, 2025

Method and system for semantic appearance transfer using splicing ViT features

Inventors: Tali Dekel (Rehovot, IL); Shai Bagon (Rehovot, IL); Omer Bar Tal (Rehovot, IL); Narek Tumanyan (Rehovot, IL)
Assignee: Yeda Research and Development Co. Ltd.
G06F3/14G06T7/11G06V10/54G06V10/56G06V10/761G06V20/70G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,696
App. No.
18/723,429
Granted
Apr 22, 2025
Kind
B2
Abstract

Using a pre-trained and fixed Vision Transformer (ViT) model as an external semantic prior, a generator is trained given only a single structure/appearance image pair as input. Given two input images, a source structure image and a target appearance image, a new image is generated by the generator in which the structure of the source image is preserved, while the visual appearance of the target image is transferred in a semantically aware manner, so that objects in the structure image are “painted” with the visual appearance of semantically related objects in the appearance image. A self-supervised, pre-trained ViT model, such as a DINO-VIT model, is leveraged as an external semantic prior, allowing for training of the generator only on a single input image pair, without any additional information (e.g., segmentation/correspondences), and without adversarial training. The method may generate high quality results in high resolution (e.g., HD).

Claims (36)

1. A method for semantic appearance transfer comprising:

capturing, by a first camera, a first image that comprises a first structure associated with a first appearance;

capturing, by a second camera, a second image that is semantically-related to the first image and is associated with a second appearance that is different from the first appearance;

providing, to a first appearance analyzer, the second image;

generating, by the first appearance analyzer, a first output that is responsive to, comprises a representation of, or comprises a function of, the second appearance;

providing, to a second appearance analyzer, the third image;

generating, by the second appearance analyzer, a second output that is responsive to, comprises a representation of, or comprises a function of, an appearance of the third image;

comparing, by a first appearance comparator, the first and second outputs;

training, a generator, to minimize or nullify the difference between the first and second outputs;

providing, to the generator, the first and second images;

processing, by the generator, using an Artificial Neural Network (ANN), the first and second images;

outputting, by the generator, based on the first and semantically-related, a third image that second images being comprises the first structure associated with the second appearance; and

displaying, by a display, the third image to a user.

2. The method according to claim 1 , further comprising correctly relating or identifying semantically matching regions or objects between the first and second images, and transferring the second appearance to the third image while maintaining high-fidelity of structure and appearance.

3. The method according to claim 1 , further comprising:

providing, to a first structure analyzer, the first image;

generating, by the first structure analyzer, a third output that is responsive to, comprises a representation of, or comprises a function of, the first structure;

providing, to a second structure analyzer, the third image;

generating, by the second structure analyzer, a fourth output that is responsive to, comprises a representation of, or comprises a function of, a structure of the third image;

comparing, by a first structure comparator, the third and fourth outputs; and

training, the generator, to minimize or nullify the difference between the third and fourth outputs.

4. The method according to claim 3 , further comprising using an optimizer with a constant learning.

5. The method according to claim 4 , wherein the optimizer comprises, uses, or is based on, an algorithm for first-order gradient-based optimization of stochastic objective functions, wherein the optimizer comprises, uses, or is based on, adaptive estimates of lower-order moments, or wherein the optimizer comprises, uses, or is based on, Adam optimizer.

6. The method according to claim 5 , further comprising using a publicly available deep learning framework that is based on, uses, or comprises, a machine learning library.

7. The method according to claim 6 , wherein the deep learning framework comprises, uses, or is based on, PyTorch.

8. The method according to claim 5 , wherein the first appearance analyzer, the second appearance analyzer, the first structure analyzer, or the second structure analyzer, comprises, uses, or is based on, a self-supervised pre-trained fixed Vision Transformer (ViT) model that uses a self-distillation approach.

9. The method according to claim 8 , wherein each of the first appearance analyzer, the second appearance analyzer, the first structure analyzer, or the second structure analyzer, comprises, uses, or is based on, a respective self-supervised pre-trained fixed Vision Transformer (ViT) model that uses a self-distillation approach.

10. The method according to claim 8 , wherein the Vision Transformer (ViT) model comprises, uses, or is based on, an architecture or code that is configured for Natural Language Processing (NLP), or wherein the Vision Transformer (ViT) model comprises, uses, or is based on, a fine-tuning or pre-trained machine learning model, wherein the Vision Transformer (ViT) model comprises, uses, or is based on, a deep learning model that comprises, is based on, or uses, attention, self-attention, or multi-head self-attention, or wherein the Vision Transformer (ViT) model further comprises splitting a respective input image into multiple non-overlapping same fixed-sized patches.

11. The method according to claim 10 , wherein the Vision Transformer (ViT) model comprises, uses, or is based on, differentially weighing a respective significance of each patch of a respective input image, wherein the Vision Transformer (ViT) model comprises, uses, or is based on, attaching or embedding a feature information to each of the patches, that comprises a key, a query, a value, a token, or any combination thereof, wherein the Vision Transformer (ViT) model comprises, uses, or is based on, attaching or embedding a position information in the image for each of the patches for providing a series of positional embedding patches, and feeding the position-embedded patches and positions to a transformer encoder, or wherein the Vision Transformer (ViT) model is further trained with image labels and is fully supervised on a dataset.

12. The method according to claim 8 , wherein the Vision Transformer (ViT) model comprises, uses, or is based on, assigning a [CLS] token to the image, predicting or estimating a class label to a respective input image, or any combination thereof.

13. The method according to claim 12 , wherein the Vision Transformer (ViT) model comprises, uses, or is based on, a DINO-ViT model.

14. The method according to claim 13 , wherein the DINO-ViT model is used as external semantic high-level prior, wherein the second appearance analyzer and the second structure analyzer are based on, or use, the same DINO-ViT, or wherein the first appearance analyzer or the second appearance analyzer comprises a DINO-ViT model.

15. The method according to claim 14 , wherein the first appearance analyzer comprises a first DINO-ViT model, and wherein the first output comprises, is a function of, uses, is based on, or is responsive to, a self-similarity of keys in the deepest attention module (Self-Sim) of the first DINO-ViT model, or wherein the second appearance analyzer comprises a second DINO-ViT model, and wherein the second output comprises, is a function of, uses, is based on, or is responsive to, a self-similarity of keys in a deepest attention module (Self-Sim) of the second DINO-ViT model.

16. The method according to claim 13 , wherein the first structure analyzer or the second structure analyzer comprises a DINO-ViT model, and wherein the first structure analyzer comprises a first DINO-VIT model, and wherein the third output comprises, is a function of, uses, is based on, or is responsive to, a [CLS] token in a deepest layer of the first DINO-ViT model, or wherein the second structure analyzer comprises a second DINO-ViT model, and wherein the fourth output comprises, is a function of, uses, is based on, or is responsive to, a [CLS] token in a deepest layer of the second DINO-ViT model.

17. The method according to claim 1 , wherein the third image comprises an across objects transfer, so that the second appearance that comprises appearance of one or more objects in the first image is transferred to semantically-related objects in the second image, while attenuating or ignoring variations in pose, number of objects, or appearance of the second image, wherein the third image comprises a within objects transfer, wherein the first appearance that comprises appearance of one or more objects in the second image is transferred between corresponding body parts or object elements of the first and second images, or wherein the second appearance comprises a global appearance information and style, color, texture, or any combination thereof, while attenuating, ignoring, or discarding, exact objects' pose, shape, perceived semantics of the objects and their surroundings, and scene's spatial layout.

18. The method according to claim 1 , wherein a [CLS] token of a DINO-ViT model applied to the second image or to the third image represents, or is responsive to, the second appearance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2024
From: DEKEL, TALI; BAR TAL, OMER; BAGON, SHAI; TUMANYAN, NAREK
To: YEDA RESEARCH AND DEVELOPMENT CO. LTD.
Reel/Frame 068175/0823 →
Continuity (2)
Provisional Application 63293842 · Dec 27, 2021
Related Publication 20240419382A1 · Dec 19, 2024
References Cited (150)
US 5092343A · Spitzer et al. · 1992 [cited by applicant]
US 5138459A · Roberts et al. · 1992 [cited by applicant]
US 5402170A · Parulski et al. · 1995 [cited by applicant]
US 5798791A · Katayama et al. · 1998 [cited by applicant]
US 6018728A · Spence et al. · 2000 [cited by applicant]
US 6038337A · Lawrence et al. · 2000 [cited by applicant]
US 6711293B1 · Lowe · 2004 [cited by applicant]
US 6897891B2 · Itsukaichi · 2005 [cited by applicant]
US 6940545B1 · Ray · 2005 [cited by applicant]
US 7432952B2 · Fukuoka · 2008 [cited by applicant]
US 8165401B2 · Funayama et al. · 2012 [cited by applicant]
US 8244053B2 · Steinberg · 2012 [cited by applicant]
US 8285067B2 · Steinberg · 2012 [cited by applicant]
US 8345984B2 · Ji et al. · 2013 [cited by applicant]
US 8705849B2 · Prokhorov · 2014 [cited by applicant]
US 8712189B2 · Bitouk et al. · 2014 [cited by applicant]
US 8773509B2 · Pan · 2014 [cited by applicant]
US 8898093B1 · Helmsen · 2014 [cited by applicant]
US 9881234B2 · Huang et al. · 2018 [cited by applicant]
US 10706335B2 · Gautam et al. · 2020 [cited by applicant]
US 20020101515A1 · Yoshida et al. · 2002 [cited by applicant]
US 20070195167A1 · Ishiyama · 2007 [cited by applicant]
US 20090102940A1 · Uchida · 2009 [cited by applicant]
US 20120249768A1 · Binder · 2012 [cited by applicant]
US 20140070613A1 · Garb et al. · 2014 [cited by applicant]
US 20140159877A1 · Huang · 2014 [cited by applicant]
US 20200160547A1 · Liu · 2020 [cited by examiner]
US 20200195910A1 · Han · 2020 [cited by examiner]
US 20200380306A1 · Hada et al. · 2020 [cited by applicant]
US 20230196712A1 · Assouline · 2023 [cited by examiner]
CN 113642585 · 2021 [cited by applicant]
JP 2006050494 · 2006 [cited by applicant]
JP 2007208922 · 2007 [cited by applicant]
JP 2008033200 · 2008 [cited by applicant]
WO 2012013914 · 2012 [cited by applicant]
WO 2020101246 · 2020 [cited by applicant]
Dmitry Ulyanov; Andrea Vedaldi; and Victor Lempitsky, entitled: “Deep Image Prior” dated May 17, 2020 [arXiv:1711.10925 [cs.CV]; arXiv:1711.10925v4 [cs.CV]; https://doi.org/10.48550/arXiv.1711.10925; https://doi.org/10.… [cited by applicant]
Hans-Petter Halvorsen, “Introduction to Database Systems”, Telemark University College, dated Mar. 3, 2014 (40 pages). [cited by applicant]
Tutorial entitled: “Oracle / SQL Tutorial” by Michael Gertz of University of California, revised Version 1.01, Jan. 2000 (66 pages). [cited by applicant]
Book published 2005 by Pearson Education, Inc. William Stallings [ISBN: 0-13-191835-4] “Wireless Communications and Networks—second Edition” (569 pages). [cited by applicant]
Telecom Regulatory Authority, “WiFi Technology”, published on Jul. 2003 (60 pages). [cited by applicant]
Bluetooth SIG published Dec. 2, 2014 standard Covered Core Package version: 4.2, “Master Table of Contents & Compliance Requirements—Specification vol. 0” (2772 pages). [cited by applicant]
Carles Gomez et al., “Overview and Evaluation of Bluetooth Low Energy: An Emerging Low-Power Wireless Technology”, published 2012 in Sensors [ISSN 1424-8220] [Sensors 2012, 12, 11734-11753; doi: 10.3390/s120211734] (20 … [cited by applicant]
Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel, “Deep ViT Features as Dense Visual Descriptors”, published Sep. 4, 2022 (arXiv:2112.05814 [cs.CV]; https://doi.org/10.48550/arXiv.2112.05814) (17 pages). [cited by applicant]
Nicholas Kolkin, Jason Salavon, and Greg Shakhnarovich entitled: “Style Transfer by Relaxed Optimal Transport and Self-Similarity”, [https://doi.org/10.48550/arXiv.1904.12785; arXiv:1904.12785v2 [cs.CV]] published 2019 … [cited by applicant]
Eli Shechtman and Michal Irani entitled: “Matching Local Self-Similarities across Images and Videos”, published 2007 in IEEE Conference on Computer Vision and Pattern Recognition [DOI: 10.1109/CVPR.2007.383198] (8 pages… [cited by applicant]
Data sheet [DS-TM4C123GH6PM-15842.2741, SPMS376E, Revision 15842.2741 Jun. 2014], “Tiva™ TM4C 123GH6PM Microcontroller—Data Sheet”, published 2015 by Texas Instruments Incorporated (1409 pages). [cited by applicant]
Narek Tumanyan; Omer Bar-Tal; Shai Bagon; and Tali Dekel, “Splicing ViT Features for Semantic Appearance Transfer”, dated Jan. 2, 2022 (arXiv:2201.00424 [cs.CV], https://doi.org/10.48550/arXiv.2201.00424) (10 pages). [cited by applicant]
Diederik P. Kingma and Jimmy Ba entitled: “Adam: A Method for Stochastic Optimization”, dated May 7, 2015 [arXiv:1412.6980 [cs.LG]; https://doi.org/10.48550/arXiv.1412.6980; Published as a conference paper at the 3rd In… [cited by applicant]
A paper by Adam Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library” dated Dec. 3, 2019 [arXiv:1912.01703 [cs.LG]; https://doi.org/10.48550/arXiv.1912.01703; Published in “Advances in Ne… [cited by applicant]
The manual “80186/80188 High-Integration 16-Bit Microprocessors” by Intel Corporation Nov. 1994 (34 pages). [cited by applicant]
JJoseph Ciaglia et al., “Absolute Beginner's Guide to Digital Photography”, published on Apr. 2004 by Que Publishing (ISBN—0-7897-3120-7) (381 pages). [cited by applicant]
Al Bovik, “Handbook of Image & Video Processing”, by Academic Press, ISBN: 0-12-119790-5, 2000 (800 pages). [cited by applicant]
Application Note No. AN1928/D (Revision 0—Feb. 20, 2001), “Roadrunner—Modular digital still camera reference design”, Freescale Semiconductor, Inc. (30 pages). [cited by applicant]
A book “The SAR Handbook: Comprehensive Methodologies for Forest Monitoring and Biomass Estimation Book”, edited by Africa Ixmucane Flores-Anderson; Kelsey E. Herndon; and Rajesh Bahadur Thapa; Emil Cherrington, publish… [cited by applicant]
Gerhard Krieger, Irena Hajnsek, and Konstantinos P. Papathanass, “A Tutorial on Synthetic Aperture Radar”, published Mar. 2013 in IEEE Geoscience and remote sensing magazine [2168-6831/13/$31.00 © 2013] entitled: (30 pa… [cited by applicant]
Adobe Digital Video Group publication, “A Digital Video Primer—An introduction to DV production, post-production, and delivery”, updated and enhanced Mar. 2004 (58 pages). [cited by applicant]
Primer by Tektronix® entitled: “A Guide to Standard and High-Definition Digital Video Measurements”, 2009 (112 pages). [cited by applicant]
IETF RFC 3640, “RTP Payload Format for Transport of MPEG-4 Elementary Streams”, Nov. 2003 (43 pages). [cited by applicant]
White paper entitled: “Understanding MPEG-4: Technologies, Advantages, and Markets—An MPEGIF White Paper”, published 2005 by The MPEG Industry Forum (Document No. mp-in-40182) (51 pages). [cited by applicant]
Gary J. Sullivan of Microsoft Corporation, “The H.264/MPEG4 Advanced Video Coding Standard and its Applications”, Standards Report published in IEEE Communications Magazine, Aug. 2006 (10 pages). [cited by applicant]
IETF RFC 3984, “RTP Payload Format for H.264 Video”, Feb. 2005 (83 pages). [cited by applicant]
Publication entitled: “An introduction to video content analysis—industry guide” published Aug. 2016 as Form No. 262 Issue 2 by British Security Industry Association (BSIA) (13 pages). [cited by applicant]
Paper entitled: “Overview of Existing Content Based Video Retrieval Systems” by Shripad A. Bhat, Omkar V. Sardessai, Preetesh P. Kunde and Sarvesh S. Shirodkar of the Department of Electronics and Telecommunication Engi… [cited by applicant]
Tinku Acharya and Ajoy K. Ray entitled: “Image Processing—Principles and Applications”, published by Wiley-Interscience [ISBN: 13-978-0-471-71998-4] (2005) (451 pages). [cited by applicant]
John G. Proakis and Dimitris G. Manolakis, “Third Edition—Digital Signal Processing—Principles, Algorithms, and Application”, published 1996 by Prentice-Hall Inc. [ISBN 0-13-394338-9] (1033 pages). [cited by applicant]
David Kriesel entitled: “A Brief Introduction to Neural Networks” (ZETA2-EN) [downloaded May 2015 from www.dkriesel.com] (244 pages). [cited by applicant]
Simon Haykin entitled: “Neural Networks and Learning Machines—Third Edition”, published 2009 by Pearson Education, Inc. [ISBN—978-0-13-147139-9] (937 pages). [cited by applicant]
Juan A. Ramirez-Quintana, Mario I. Cacon-Murguia, and F. Chacon-Hinojos entitled: “Artificial Neural Image Processing Applications: A Survey”, Engineering Letters, 20:1, EL_20_1_09 (Advance online publication: Feb. 27, … [cited by applicant]
M. Egmont-Petersen, D. de Ridder, and H. Handels, “Image processing with neural networks—a review”, by Pattern Recognition Society in Pattern Recognition 35 (2002) 2279-2301 (23 pages). [cited by applicant]
Christian Szegedy, Alexander Toshev, and Dumitru Erhan, “Deep Neural Networks for Object Detection”, of Google, Inc. (downloaded Jul. 2015) (9 pages). [cited by applicant]
Dumitru Erhan, Christian Szegedy, Alexander Toshev, and Dragomir Anguelov (of Google, Inc., Mountain-View, California, U.S.A.), “Scalable Object Detection using Deep Neural Networks”, a CVPR2014 paper provided by the Co… [cited by applicant]
Shawn McCann and Jim Reesman (both of Stanford University), “Object Detection using Convolutional Neural Networks” (downloaded Jul. 2015) (5 pages). [cited by applicant]
Mehdi Ebady Manaa, Nawfal Turki Obies, and Dr. Tawfiq A. Al-Assadi (of Department of Computer Science, Babylon University), “Object Classification using neural networks with Gray-level Co-occurrence Matrices (GLCM)”, (d… [cited by applicant]
Dan C. Ciresan et al., “High-Performance Neural Networks for Visual Object Classification”, a technical report No. IDSIA-01-11 Jan. 2011 published by IDSIA/USI-SUPSI (12 pages). [cited by applicant]
Yuhua Zheng et al. “Object Recognition using Neural Networks with Bottom-Up and top-Down Pathways” (downloaded Jul. 2015) (13 pages). [cited by applicant]
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman (all of Visual Geometry Group, University of Oxford), “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps” Apr. 19, 2014, (… [cited by applicant]
User's Guide Version 4 by Howard Demuth and Mark Beal, “Neural Network ToolBox—For Use with MATLAB®”, MathWorks, Inc. (Headquartered in Natick, MA, U.S.A.) published Jul. 2002 (840 pages). [cited by applicant]
A book entitled: “Introduction to Deep Learning From Logical Calculus to Artificial Intelligence” by Sandro Skansi [ISSN 1863-7310 ISSN 2197-1781, ISBN 978-3-319-73003-5], published 2018 by Springer International Publis… [cited by applicant]
Weibo Liua, Zidong Wanga, Xiaohui Liua, Nianyin Zengb, Yurong Liuc, and Fuad E. Alsaadid, “A Survey of Deep Neural Network Architectures and Their Applications”, published Dec. 2016 [DOI: 10.1016/j.neucom.2016.12.038] (… [cited by applicant]
Dilip K. Prasad, “Survey of The Problem of Object Detection In Real Images” published International Journal of Image Processing (IJIP), vol. 6, Issue 6—2012 (26 pages). [cited by applicant]
A. Ashbrook and N. A. Thacker, “Tutorial: Algorithms For 2-dimensional Object Recognition”, published by the Imaging Science and Biomedical Engineering Division of the University of Manchester, Tina Memo No. 1996-003, J… [cited by applicant]
Computer Vision: Mar. 2000 Chapter 4 entitled: “Pattern Recognition Concepts” (38 pages). [cited by applicant]
Isabelle Guyon, et al., “Hands-On Pattern Recognition—Challenges in Machine Learning, vol. 1”, published by Microtome Publishing, 2011 (ISBN-13:978-0-9719777-1-6) (482 pages). [cited by applicant]
Paul Viola and Michael Jones, “Robust Real-Time Face Detection”, International Journal of Computer Vision 2004 pp. 137-154 (18 pages). [cited by applicant]
Paul Viola and Michael Jones, “Rapid Object Detection using a Boosted Cascade of Simple Features”, Accepted Conference on Computer Vision and Pattern Recognition 2001 (9 pages). [cited by applicant]
Technical report No. RL-TR-94-150 by Rome Laboratory, Air force Material Command, Griffiss Air Force Base, New York, entitled: “Neural Network Communications Signal Processing”, published Aug. 1994 (128 pages). [cited by applicant]
Image-net.org/download-API, 2014 Stanford Vision Lab (3 pages). [cited by applicant]
Fei-Fei Li and Olga Russakovsky (ICCV 2013), “Analysis of large Scale Visual Recognition”, (60 pages). [cited by applicant]
ImageNet presentation by Fei-Fei Li (of Computer Science Dept., Stanford University), “Crowdsourcing, benchmarking, & other cool things”, 2010 Stanford Vision Lab (64 pages). [cited by applicant]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton (all of University of Toronto), “ImageNet Classification with Deep Convolutional Neural Networks”, May 18, 2015 (35 pages). [cited by applicant]
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi, “You Only Look Once: Unified, Real-Time Object Detection”, published May 9, 2016 (10 pages). [cited by applicant]
Juan Du of New Research and Development Center of Hisense, Qingdao 266071, China, “Understanding of Object Detection Based on CNN Family and YOLO”, published 2018 in IOP Conf. Series: Journal of Physics: Conf. Series 10… [cited by applicant]
Joseph Redmon and Ali Farhadi, “YOLO9000: Better, Faster, Stronger”, published 2016 (9 pages). [cited by applicant]
Duy Thanh Nguyen, Tuan Nghia Nguyen, Hyun Kim, and Hyuk-Jae Lee, “A High-Throughput and Power-Efficient FPGA Implementation of YOLO CNN for Object Detection”, published Aug. 2019 [DOI: 10.1109/TVLSI.2019.2905242] in IEE… [cited by applicant]
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation”, published 2014 In Proc. IEEE Conf. on computer vision and pattern reco… [cited by applicant]
Ross Girshick of Microsoft Research, “Fast R-CNN”, published Sep. 27, 2015 [arXiv:1504.08083v2 [cs. CV]] In Proc. IEEE Intl. Conf. on computer vision, pp. 1440-1448. 2015 (9 pages). [cited by applicant]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal networks”, published 2015 (9 pages). [cited by applicant]
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar, “Focal Loss for Dense Object Detection”, published Feb. 7, 2018 in IEEE Transactions on Pattern Analysis and Machine Intelligence. 42 (2): 318-327 … [cited by applicant]
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie, “Feature Pyramid Networks for Object Detection”, published Apr. 19, 2017 [arXiv:1612.03144v2 [cs.CV]] (10 pages). [cited by applicant]
Mxing Li and Fengbo Ren, “Light-Weight RetinaNet for Object Detection”, published May 24, 2019 [arXiv: 1905.10011v1 [cs.CV]] (10 pages). [cited by applicant]
Jie Zhou, Ganqu Cui, Shengding Hu, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun, “Graph neural networks: A review of methods and applications”, published at AI Open 2021 [arXiv:18… [cited by applicant]
Isaac Ronald Ward, Jack Joyner, Casey Lickfold, Stash Rowe, Yulan Guo, and Mohammed Bennamoun, “A Practical Guide to Graph Neural Networks”, published 2020 [arXiv:2010.05234 [cs.LG]] (28 pages). [cited by applicant]
Rex Ying, Yuanfang Li, and Xin Li of Stanford University, “GraphNet: Recommendation system based on language and network structure”, published 2017 by Stanford University (9 pages). [cited by applicant]
Anton Tsitsulin, John Palowitch, Bryan Perozzi, and Emmanuel Müller, “Graph Clustering with Graph Neural Networks”, published Jun. 30, 2020 [arXiv:2006.16904v1 [cs.LG]] (13 pages). [cited by applicant]
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam of Google Inc., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applicat… [cited by applicant]
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen of Google Inc., “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, published Mar. 21, 2019 [arXiv:1801.04381v4 [cs.CV]] (14 pages). [cited by applicant]
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam, “Searching for MobileNetV3”, published 2019 [arXiv:19… [cited by applicant]
Jonathan Long, Evan Shelhamer, and Trevor Darrell, “Fully Convolutional Networks for Semantic Segmentation”, published Apr. 1, 2017 in IEEE Transactions on Pattern Analysis and Machine Intelligence ( vol. 39, Issue: 4) … [cited by applicant]
Hongyang Gao and Shuiwang Ji, “Graph U-Nets”, published 2019 [arXiv:1905.05178 [cs.LG]] (10 pages). [cited by applicant]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation”, published May 18, 2015 in Medical Image Computing and Computer-Assisted Intervention (MICCAI), Sprin… [cited by applicant]
Simonyan and Zisserman from Visual Geometry Group (VGG) at University of Oxford, “Very Deep Convolutional Networks for Large-Scale Image Recognition”, published 2015 [arXiv:1409.1556 [cs.CV]] as a conference paper at IC… [cited by applicant]
“VGG16—Convolutional Network for Classification and Detection”, published Nov. 20, 2018 in ‘Popular networks’ (27 pages). [cited by applicant]
David G. Lowe of the Keypoints Computer Science Department University of British Columbia Vancouver, B.C., Canada, entitled: “Distinctive Image Features from Scale-Invariant”, published Jan. 5, 2004 [International Journ… [cited by applicant]
Tony Lindeberg entitled: “Scale Invariant Feature Transform”, published May 2012 [DOI: 10.4249/scholarpedia.10491] (18 pages). [cited by applicant]
Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool, all of ETH Zurich, entitled: “SURF: Speeded Up Robust Features”, presented at the ECCV 2006 conference and published 2008 at Computer Vision and Image Unders… [cited by applicant]
Edward Rosten and Tom Drummond of the Department of Engineering, Cambridge University, UK, “Machine Learning for High-speed Corner Detection”, published 2006 in Computer Vision—ECCV 2006 [Lecture Notes in Computer Scien… [cited by applicant]
Article published in IEEE Transactions on Pattern Analysis and Machine Intelligence (vol. 32, Issue: 1, Jan. 2010) [DOI: 10.1109/TPAMI.2008.275] entitled: “FASTER and better: A Machine Learning Approach to Corner Detect… [cited by applicant]
DocID022930 Rev. 6 dated Apr. 2015 entitled: “SPBT2632C1A—Bluetooth® technology class-1 module” (27 pages). [cited by applicant]
Chapter 20: “Wireless Technologies” of the publication No. 1-587005-001-3 by Cisco Systems, Inc. (7/99) “Internetworking Technologies Handbook” (42 pages). [cited by applicant]
IPhone 6 technical specification (retrieved Oct. 2015 from www.apple.com/iphone-6/specs/) (32 pages). [cited by applicant]
User Guide, “iPhone User Guide For iOS 8.4 Software”, dated 2015 (019-00155/2015-06) by Apple Inc. (196 pages). [cited by applicant]
User manual numbered English (EU), “SM-G925F SM-G925FQ SM-G925I User Manual” Mar. 2015 (Rev. 1.0) (145 pages). [cited by applicant]
Galaxy S6 Edge—Technical Specification (retrieved Oct. 2015 from www.samsung.com/us/explore/galaxy-s-6-features-and-specs) (1 page). [cited by applicant]
“Android Tutorial”, downloaded from tutorialspoint.com on Jul. 2014 (216 pages). [cited by applicant]
“IOS Tutorial”, downloaded from tutorialspoint.com on Jul. 2014 (185 pages). [cited by applicant]
Muhammad Ayoub Kamal, Hafiz Wahab Raza, Muhammad Mansoor Alam, and Mazliham Mohd Su'ud, “Highlight the Features of AWS, GCP and Microsoft Azure that Have an Impact when Choosing a Cloud Service Provider”, ‘International… [cited by applicant]
Dac-Nhuong Le et al., “Cloud Computing and Virtualization”, published 2018 by John Wiley & Sons, Inc. [ISBN 978-1-119-48790-6] (223 pages). [cited by applicant]
EBook authored by Gustavo Alessandro Andrade Santana, “Data Center Virtualization Fundamentals”, published 2014 by Cisco Systems, Inc. (Cisco Press) [ISBN-13: 978-1-58714-324-3] (1406 pages). [cited by applicant]
IBM RedBook entitled: “IBM PowerVM Virtualization—Introduction and Configuration” published by IBM Corporation Jun. 2013 (790 pages). [cited by applicant]
IBM Corporation, “Power Systems—Introduction to virtualization”, published 2009 (54 pages). [cited by applicant]
Microsoft publication entitled: “Inside-Out Windows Server 2012”, by William R. Stanek, published 2013 by Microsoft Press (800 pages). [cited by applicant]
“UNIX Tutorial” by tutorialspoint.com, downloaded on Jul. 2014 (152 pages). [cited by applicant]
“Windows Internals—Part 1” by Mark Russinovich, David A. Solomon, and Alex Ioescu, published by Microsoft Press in 2012 (752 pages). [cited by applicant]
“Windows Internals—Part 2”, by Mark Russinovich, David A. Solomon, and Alex Ioescu, published by Microsoft Press in 2012 (672 pages). [cited by applicant]
Jerry Honeycutt, “Introducing Windows 8—An Overview for IT Professionals” Microsoft Press 2012 (168 pages). [cited by applicant]
Samsung Electronics Co., Ltd. presentation entitled: “Google™ Chrome OS User Guide” published 2011 (41 pages). [cited by applicant]
Application Note No. RES05B00008-0100/Rec. 1.00 by Renesas Technology Corp. entitled: “R8C Family—General RTOS Concepts”, published Jan. 2010 (20 pages). [cited by applicant]
Walter Cedeno et al, “An Overview of Real-Time Operating Systems”, The Association for Laboratory Automation, Feb. 2007 (6 pages). [cited by applicant]
Chapter 2 entitled: “Basic Concepts of Real Time Operating Systems” of a book entitled: “Hardware-Dependent Software—Principles and Practice”, published 2009 [ISBN—978-1-4020-9435-4] by Springer Science + Business Media… [cited by applicant]
Nicolas Melot, “Study of an operating system: FreeRTOS—Operating systems for embedded devices”, (downloaded Jul. 2015) (39 pages). [cited by applicant]
Dr. Richard Wall entitled: “Carebot PIC32 MX7ck implementation of Free RTOS”, (dated Sep. 23, 2013) (18 pages). [cited by applicant]
“FreeRTOS™ Modules” published in the www,freertos.org web-site dated Nov. 26, 2006 (112 pages). [cited by applicant]
Rich Goyette of Carleton University as part of ‘SYSC5701: Operating System Methods for Real-Time Applications’, entitled: “An Analysis and Description of the Inner Workings of the FreeRTOS Kernel”, published Apr. 1, 200… [cited by applicant]
Ashish Vaswani; Noam Shazeer; Niki Parmar; Jakob Uszkoreit; Llion Jones; Aidan N. Gomez; Lukasz Kaiser; and Illia Polosukhin, “Attention Is All You Need”, dated Jun. 12, 2017 (arXiv:1706.03762 [cs.CL]; https://doi.org/1… [cited by applicant]
Alexey Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, dated Oct. 22, 2020 (arXiv:2010.11929 [cs.CV]; https://doi.org/10.48550/arXiv.2010.11929) (22 pages). [cited by applicant]
Mathilde Caron; Hugo Touvron; Ishan Misra; Hervé Jégou; Julien Mairal; Piotr Bojanowski; and Armand Joulin, “Emerging Properties in Self-Supervised Vision Transformers”, dated Apr. 29, 2021 (arXiv:2104.14294 [cs.CV]) (2… [cited by applicant]
International Search Report of PCT/IL2022/051346 dated Aug. 17, 2023. [cited by applicant]
Written Opinion of PCT/IL2022/051346 dated Aug. 17, 2023. [cited by applicant]
Park, S. et al., “Neural Crossbreed: Neutral Based Image Metamorphosis”, ACM Transactions on Graphics, Publication (online) Sep. 2, 2020 (16 pages). [cited by applicant]