IP Library Granted Patent US 12675995
Granted Patent B2
US 12675995 · App. 18/519,102 · Granted Jul 7, 2026

Multimedia image processing method, electronic device, terminal device connected thereto, and non-transitory computer-readable recording medium

Inventors: Pin-Yu Chou (Taipei, TW); Yueh-Hua Lee (Taipei, TW); Ming-Hsien Wu (Taipei, TW); Hui-Mei Hung (Taipei, TW)
Assignee: COMPAL ELECTRONICS, INC.
G06V20/41G06T7/70G06V10/77G06V40/174G06V40/20G06T2207/10016G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12675995
App. No.
18/519,102
Granted
Jul 7, 2026
Kind
B2
Abstract

A method for multimedia image processing includes steps of identifying objects and selecting images, and is characterized in that when a body position of a preset object is detected and conformed to match a preset posture; or a plurality of the preset objects are detected to have body movements and facial expressions that are emotional, and the preset objects have at least one of looking in similar directions, one looking at the other, and at least two looking at each other, an interception time point is selected for selecting a candidate image, and the candidate image can be collected to produce a concatenated video with rich contents. An electronic device is also introduced for multimedia image processing, a terminal device connected thereto, and a non-transitory computer-readable recording medium.

Claims (33)

1 . A multimedia image processing method, executed by an electronic device reading an executable code, identifying a plurality of preset objects by artificial intelligence, and performing multimedia image processing to the plurality of preset objects, comprising the following steps:

identifying objects: from an initial image, identifying whether there are the plurality of preset objects by artificial intelligence, and detecting bodies and facial expressions of the plurality of preset objects in the initial image; and

selecting images: setting a selection condition, the selection condition comprises the plurality of preset objects having body movements and facial expressions that are emotional, and the plurality of preset objects have at least one of looking in similar directions, one looking at the other, and at least two looking at each other, an interception time point is selected in the initial image when conforming to the selection condition, and a candidate image is selected according to the interception time point in the initial image, the candidate image is configured to be collected to produce a concatenated video.

2 . The multimedia image processing method according to claim 1 , wherein at least one of the preset objects is a child, and at least one is an adult; the selection condition further comprises calculating that the facial expression of the child in the initial image is a positive emotion or negative emotion, and the facial expression of the adult in the initial image is a positive emotion.

3 . The multimedia image processing method according to claim 2 , wherein the selection condition further comprises calculating that the movements of the child and the adult in the initial image are the same or opposite, and/or the expressions are the same or opposite.

4 . The multimedia image processing method according to claim 1 , wherein the selection condition comprises detecting that there is a selected article in the initial image, and when at least one of the plurality of preset objects is looking at the selected article, and/or at least one is holding the selected article the interception time point is selected in the initial image.

5 . The multimedia image processing method according to claim 1 ,

wherein the plurality of preset objects respectively define a positioning frame in the initial image, each of which is an area occupied by the corresponding preset object in the initial image, the selection condition comprises detecting that when an area overlapping a plurality of positioning frames in the initial image is greater than a threshold, the interception time point is selected in the initial image.

6 . The multimedia image processing method according to claim 1 , wherein a neutral axis is defined in the initial image, and the plurality of preset objects respectively define a human body posture, the selection condition comprises determining that when the human body posture of each preset object in the initial image deviates from the neutral axis by a threshold the interception time point is selected in the initial image.

7 . The multimedia image processing method according to claim 1 , wherein the plurality of preset objects respectively define a body position, the selection condition comprises determining the body position of each preset object in the initial image, and determining that when the angle of arms opening of each preset object is greater than a threshold the interception time point is selected in the initial image.

8 . A terminal device in communication with the electronic device executing the method according to claim 1 , the terminal device is equipped with an application program, the terminal device executes the application program to collect more than one candidate image to produce the concatenated video.

9 . A multimedia image processing method, executed by an electronic device reading an executable code, identifying a preset object by artificial intelligence, and performing multimedia image processing to the preset object, comprising the following steps:

identifying objects: from an initial image, identifying whether there is the preset object by artificial intelligence, detecting a body position of the preset object in the initial image; and

selecting images: setting a selection condition, the selection condition comprises the body position of the preset object conforming to a preset posture, an interception time point is selected in the initial image when conforming to the selection condition, and a candidate image is selected according to the interception time point in the initial image, the candidate image is configured to be collected to produce a concatenated video,

wherein a neutral axis is defined in the initial image, and the preset object defines a human body posture, the selection condition comprises detecting that when the human body posture of the preset object in the initial image deviates from the neutral axis by a threshold, the interception time point is selected in the initial image.

10 . The multimedia image processing method according to claim 9 , wherein the selection condition comprises detecting the body position of the preset object in the initial image, and determining that when the angle of arms opening of the preset object is greater than a threshold the interception time point is selected in the initial image.

11 . The multimedia image processing method according to claim 9 , wherein the preset object defines a human body midline in the initial image, the selection condition comprises detecting the body position of the preset object in the initial image, and determining that when the body position of the preset object has body stretching, body connection, body constituting geometry, and/or body symmetry, the interception time point is selected in the initial image.

12 . A terminal device in communication with the electronic device executing the method according to claim 9 , the terminal device is equipped with an application program, the terminal device executes the application program to collect more than one candidate image to produce the concatenated video.

13 . An electronic device for multimedia image processing, comprising:

a photographic unit for taking an initial image;

an intelligent identification unit, electrically connected to the photographic unit to receive the initial image, identifying an initial image having a plurality of preset objects by artificial intelligence, and detecting bodies and facial expressions of the plurality of preset objects in the initial image; and

an intelligent processing unit, the intelligent processing unit is electrically connected to the intelligent identification unit, and the intelligent processing unit reads an executable code and executes multimedia image processing to the plurality of preset objects, comprising setting a selection condition, the selection condition comprises the plurality of preset objects have body movements and facial expressions that are emotional, and the plurality of preset objects have at least one of looking in similar directions, one looking at the other, and at least two looking at each other, an interception time point is selected in the initial image when conforming to the selection condition, the candidate image is configured to be collected to produce a concatenated video.

14 . The electronic device for multimedia image processing according to claim 13 , wherein the photographic unit and the intelligent identification unit belong to a physical host, and the intelligent processing unit belongs to a cloud host.

15 . The electronic device for multimedia image processing according to claim 13 , wherein the preset object comprises at least one child and one adult, the intelligent identification unit further comprises an expression identification module, configured to identify the expressions of the child and the adult; a body identification module, configured to identify the body position of the child and the adult; a viewing angle identification module, configured used to identify viewing angles of the child and the adult; and/or a specific article identification module, configured to identify a specific article in the initial image.

16 . A terminal device in communication with the electronic device according to claim 13 , the terminal device is equipped with an application program, the terminal device executes the application program to collect more than one candidate image to produce the concatenated video.

17 . An electronic device for multimedia image processing, comprising:

a photographic unit for taking an initial image;

an intelligent identification unit, electrically connected to the photographic unit to receive the initial image, identifying the initial image having a preset object by artificial intelligence, and detecting a body position of the preset object; and

an intelligent processing unit, electrically connected to the intelligent identification unit, and reading an executable code and executing thereto, in order to set a selection condition, the selection condition comprises the body position of the preset object conforming to a preset posture, an interception time point is selected in the initial image when conforming to the selection condition, the intelligent processing unit selects a candidate image according to the interception time point, the candidate image is configured to be collected to produce a concatenated video,

wherein a neutral axis is defined in the initial image, and the preset object defines a human body posture, the selection condition comprises detecting that when the human body posture of the preset object in the initial image deviates from the neutral axis by a threshold, the interception time point is selected in the initial image.

18 . The electronic device for multimedia image processing according to claim 17 , wherein the photographic unit and the intelligent identification unit belong to a physical host, and the intelligent processing unit belongs to a cloud host.

19 . The electronic device for multimedia image processing according to claim 17 , wherein the preset object comprises at least one child and one adult, the intelligent identification unit further comprises an expression identification module, configured to identify the expressions of the child and the adult; a body identification module, configured to identify the body position of the child and the adult; a viewing angle identification module, configured to identify viewing angles of the child and the adult; and/or a specific article identification module, configured to identify a specific article in the initial image.

20 . A terminal device in communication with the electronic device according to claim 17 , the terminal device is equipped with an application program, the terminal device executes the application program to collect more than one candidate image to produce the concatenated video.