DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06 July 2026 has been entered.
Priority
Acknowledgment is made of applicant's claim for foreign priority based on an application filed in India on 01/12/2024. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Response to Arguments
Applicant’s arguments with respect to claims 1-2, 8-13, and 16-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 8-10, 12, 13, 16-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lim et al. (U.S. Patent Application Publication 2020/0210766) in view of Kottur et al. (U.S. Patent 12,142,298) in view of Park (KR20180025367A).
Regarding claim 1, Lim et al. discloses an aspect ratio based method for displaying a video, the aspect ratio based method comprising: identifying a primary region, within each of video frames of a video, based on an analysis of the video frames (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); determining a first aspect ratio in which the video is displayed on at least one of an electronic device or one or more applications (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); predicting, using an Artificial Intelligence (AI) model, positions of the primary region, based on the determined first aspect ratio (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); obtaining frames matching the determined first aspect ratio and having the predicted positions of the primary region (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); and displaying the video using the obtained frames at the determined first aspect ratio (paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame). However, Lim et al. fails to disclose obtaining a multi-modal contextual input from a user input; identifying a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of a plurality of video frames of the video, based on an analysis of the plurality of video frames and the multi-modal contextual input, wherein the one or more secondary events are contextually related to the primary event; and obtaining a semantic relationship between identified primary event and the one or more secondary events.
Referring to the Kottur et al. reference, Kottur et al. discloses a method of displaying a video, the method comprising: obtaining a multi-modal contextual input from a user input (Fig. 5; col. 1, line 66 – col. 2, line 17 - the assistant system may enable the user to interact with the assistant system via user inputs of various modalities (e.g., audio, voice, text, image, video, gesture, motion, location, orientation) in stateful and multi-turn conversations to receive assistance from the assistant system – multi-modal inputs (e.g., voice inputs and text inputs), hybrid/multi-modal inputs, or any combination thereof – user inputs provided by a user may be associated with particular assistant-related tasks, and may include, for example, user requests (e.g., verbal requests for information or performance of an action), user interactions with an assistant application associated with the assistant system (e.g., selection of UI elements via touch or gesture), or any other type of suitable user input that may be detected and understood by the assistant system (e.g., user movements detected by the client device of the user); col. 46, lines - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot); and identifying events in, within each of a plurality of video frames of the video, based on analysis of the plurality of video frames and the multi-modal contextual input (Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had obtained a multi-modal contextual input from a user input and identified events in, within each of a plurality of video frames of the video, based on analysis of the plurality of video frames and the multi-modal contextual input as disclosed by Kottur et al. in the method disclosed by Lim et al. in order to allow a user to create a video story with highlights and edit clips hands-free, using natural language. However, Lim et al. in view of Kottur et al. still fails to disclose identifying a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event; and obtaining a semantic relationship between identified primary event and the one or more secondary events.
Referring to the Park reference, Park discloses a method for displaying a video, the method comprising: identifying a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event (paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data); and obtaining a semantic relationship between identified primary event and the one or more secondary events (paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had identified a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event; and obtained a semantic relationship between identified primary event and the one or more secondary events as disclosed by Park in the method disclosed by Lim et al. in view of Kottur et al. in order to ensure that the regions of interest are further enhanced. Once Lim et al., Kottur et al., and Park are combined, the obtained frames matching the determined first aspect ratio disclosed by Lim et al. would include the primary event and the one or more secondary events located at the predicted positions in the obtained frames.
Regarding claim 2, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 1 including that the aspect ratio based method further comprises identifying the primary event and the one or more secondary events, based on an analysis of at least one of an audio of the plurality of video frames, the plurality of video frames, or a plurality of multi-modal contextual inputs (Lim et al.: paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; Kottur et al: Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features; Park: paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow).
Regarding claim 8, Lim et al. in the Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 1 including that wherein the obtaining the semantic relationship between the identified primary event and the one or more secondary events comprises: identifying at least one of one or more objects, one or more faces, an orientation of a head of one or more users, or gaze angles of the one or more users within the primary event and the one or more secondary events based on the analysis of the plurality of video frames; and obtaining, based on the identifying the at least one of the one or more objects, the one or more faces, the orientation of the head of the one or more users, or the gaze angles of the one or more users identified within the primary region and the one or more secondary regions, the semantic relationship between the primary event and each of the one or more secondary events based on a plurality of semantic relationship parameters, wherein the plurality of semantic relationship parameters comprise proximity, relative to a camera, of the one or more objects or the one or more faces, the gaze angles of the one or more users, a pixel displacement in the primary region and the one or more secondary regions, a visual similarity in the primary event, and a visual similarity in the one or more secondary regions (Lim et al.: paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; Kottur et al.: Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features; Park: paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data).
Regarding claim 9, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 1 including that wherein the first aspect ratio is determined based on a second aspect ratio of at least one of a display of the electronic device or the one or more applications, wherein the display is configured to display the video (Lim et al.: paragraph [0005] – provided are an image processing apparatus that is capable of acquiring an output image while minimizing image distortion through detection of an area of an area of interest and adjusting an aspect ratio of an input image, and an image processing method thereof; paragraph [0087] – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame).
Regarding claim 10, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 1 including that wherein obtaining the frames comprises: obtaining background features of the primary event and each of the one or more secondary events based on the semantic relationship and an assigned second priority; determining a plurality of aesthetic effects for the primary event and each of the one or more secondary events based on the background features and an event score from the semantic relationship; and obtaining the frames matching the determined first aspect ratio, wherein the primary event and the one or more secondary events are located at the predicted positions, and wherein the obtained frames comprise the determined plurality of aesthetic effects (Lim et al.: paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; Kottur et al.: Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features; Park: paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0028] – the streaming service apparatus 200 may not only divide the execution screen of the content into the main area and the auxiliary area but also divide the main figure in stages and set the encoding quality in stages according to the steps of the main figure; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data; plurality of aesthetic effects – aspect ratio and quality/resolution).
Regarding claim 12, Lim et al. discloses an electronic device for displaying a video, the electronic device comprising: memory storing instructions; one or more processors, wherein the instructions, when executed by the one or more processors individually or collectively, cause the electronic device to: identify a primary region within each video frames of a video based on an analysis of the video frames (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); determine a first aspect ratio in which the video is to be displayed on at least one of an electronic device or one or more applications (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); predict, using an Artificial Intelligence (AI) model, positions of the primary region, based on the determined first aspect ratio (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); obtain frames matching the determined first aspect ratio and having the predicted positions of the primary region (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); and display the video using the obtained frames at the determined first aspect ratio (paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame). However, Lim et al. fails to disclose obtain a multi-modal contextual input from a user input; identify a primary event in a primary region and one or more secondary events in one or more secondary regions within each of a plurality of video frames of the video based on an analysis of the plurality of video frames and the multi-modal contextual input, wherein the one or more secondary events are contextually related to the primary event; and obtaining a semantic relationship between identified primary event and the one or more secondary events.
Referring to the Kottur et al. reference, Kottur et al. discloses an electronic device for displaying video, the electronic device comprising: obtaining a multi-modal contextual input from a user input (Fig. 5; col. 1, line 66 – col. 2, line 17 - the assistant system may enable the user to interact with the assistant system via user inputs of various modalities (e.g., audio, voice, text, image, video, gesture, motion, location, orientation) in stateful and multi-turn conversations to receive assistance from the assistant system – multi-modal inputs (e.g., voice inputs and text inputs), hybrid/multi-modal inputs, or any combination thereof – user inputs provided by a user may be associated with particular assistant-related tasks, and may include, for example, user requests (e.g., verbal requests for information or performance of an action), user interactions with an assistant application associated with the assistant system (e.g., selection of UI elements via touch or gesture), or any other type of suitable user input that may be detected and understood by the assistant system (e.g., user movements detected by the client device of the user); col. 46, lines - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot); and identifying events in, within each of a plurality of video frames of the video, based on analysis of the plurality of video frames and the multi-modal contextual input (Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had obtained a multi-modal contextual input from a user input and identified events in, within each of a plurality of video frames of the video, based on analysis of the plurality of video frames and the multi-modal contextual input as disclosed by Kottur et al. in the device disclosed by Lim et al. in order to allow a user to create a video story with highlights and edit clips hands-free, using natural language. However, Lim et al. in view of Kottur et al. still fails to disclose identifying a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event; and obtaining a semantic relationship between identified primary event and the one or more secondary events.
Referring to the Park reference, Park discloses an electronic device for displaying a video, the electronic device comprising one or more processors configured to: identifying a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event (paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data); and obtaining a semantic relationship between identified primary event and the one or more secondary events (paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had identified a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event; and obtained a semantic relationship between identified primary event and the one or more secondary events as disclosed by Park in the device disclosed by Lim et al. in view of Kottur et al. in order to ensure that the regions of interest are further enhanced. Once Lim et al., Kottur et al., and Park are combined, the obtained frames matching the determined first aspect ratio disclosed by Lim et al. would include the primary event and the one or more secondary events located at the predicted positions in the obtained frames.
Regarding claim 13, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 12 including that wherein the primary event and the one or more secondary events are identified based on an analysis of at least one of an audio of the plurality of video frames, the plurality of video frames or a plurality of multi-modal contextual inputs (Lim et al.: paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; Kottur et al.: Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features; Park: paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user’s character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow).
Regarding claim 16, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 12 including that wherein the instructions, when executed by the one or more processors individually or collectively, cause the electronic device to: identify at least one of one or more objects, one or more faces, orientation of a head of one or more users, or gaze angles of the one or more users within the primary event and the one or more secondary events based on the analysis of the plurality of video frames; and obtain, based on the at least one of the one or more objects, the one or more faces, the orientation of the head of the one or more users, or the gaze angles of the one or more users identified within the primary region and the one or more second regions, the semantic relationship between the primary event and each of the one or more secondary events based on a plurality of semantic relationship parameters, wherein the plurality of semantic relationship parameters comprise proximity, relative to a camera, of the one or more objects or the one or more faces, the gaze angles of the one or more users, a pixel displacement in the primary region and the one or more secondary regions, a visual similarity in the primary event, and a visual similarity in the one or more secondary regions (Lim et al.: paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; Kottur et al.: Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features; Park: paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data).
Regarding claim 17, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 12 including that the electronic device, further comprises a display configured to display the video, wherein the first aspect ratio is determined based on a second aspect ratio of at least one of a display or the one or more applications (Lim et al.: paragraph [0005] – provided are an image processing apparatus that is capable of acquiring an output image while minimizing image distortion through detection of an area of an area of interest and adjusting an aspect ratio of an input image, and an image processing method thereof; paragraph [0087] – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame).
Regarding claim 18, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claim 12 including that wherein the instructions, when executed by the one or more processors individually or collectively, cause the electronic device to: obtain background features of the primary event and each of the one or more secondary events based on the semantic relationship and an assigned second priority; determine a plurality of aesthetic effects for the primary event and each of the one or more secondary events based on the background features and an event score from the semantic relationship; and obtain the frames matching the determined first aspect ratio, wherein the primary event and the one or more secondary events are located at the predicted positions in the obtained frames, and wherein the obtained frames comprise the determined plurality of aesthetic effects (Lim et al.: paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; Kottur et al.: Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features; Park: paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0028] – the streaming service apparatus 200 may not only divide the execution screen of the content into the main area and the auxiliary area but also divide the main figure in stages and set the encoding quality in stages according to the steps of the main figure; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data; plurality of aesthetic effects – aspect ratio and quality/resolution).
Regarding claim 20, Lim et al. discloses a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: identify a primary region, within each of video frames of a video, based on an analysis of the video frames (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); determine a first aspect ratio in which the video is displayed on at least one of an electronic device or one or more applications (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); predict, using an Artificial Intelligence (AI) model, positions of the primary region, based on the determined first aspect ratio (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); obtain frames matching the determined first aspect ratio and having the predicted positions of the primary region (paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame); and display the video using the obtained frames at the determined first aspect ratio (paragraph [0057] – the memory 110 stores an instruction that controls the processor 120 to acquire an output image frame based on information on an area of interest acquired by applying an input image frame on a learning network model – here, the learning network model may be a model that is trained to acquire information on an area of interest in an input image frame; paragraph [0087] – when information on an area of interest is acquired from a learning network model, the processor 120 may retarget an input image frame by applying the first conversion weight to pixels corresponding to the area of interest, and applying the second conversion weight to pixels corresponding to the remaining area (or an area of non-interest) – the retargeting information may include the aspect ratio of the input image frame and the aspect ratio of the output image frame). However, Lim et al. fails to disclose obtain a multi-modal contextual input from a user input; identify a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of a plurality of video frames of a video, based on an analysis of the plurality of video frames and the multi-modal contextual input; and obtaining a semantic relationship between identified primary event and the one or more secondary events.
Referring to the Kottur et al. reference, Kottur et al. discloses a method of displaying a video, the method comprising: obtaining a multi-modal contextual input from a user input (Fig. 5; col. 1, line 66 – col. 2, line 17 - the assistant system may enable the user to interact with the assistant system via user inputs of various modalities (e.g., audio, voice, text, image, video, gesture, motion, location, orientation) in stateful and multi-turn conversations to receive assistance from the assistant system – multi-modal inputs (e.g., voice inputs and text inputs), hybrid/multi-modal inputs, or any combination thereof – user inputs provided by a user may be associated with particular assistant-related tasks, and may include, for example, user requests (e.g., verbal requests for information or performance of an action), user interactions with an assistant application associated with the assistant system (e.g., selection of UI elements via touch or gesture), or any other type of suitable user input that may be detected and understood by the assistant system (e.g., user movements detected by the client device of the user); col. 46, lines - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot); and identifying events in, within each of a plurality of video frames of the video, based on analysis of the plurality of video frames and the multi-modal contextual input (Fig. 5; col. 3, lines 14-33 – a digital story may be a way for the user to relive their memories, which could be connected by time or semantic meanings - a digital story may be in a video format with clips, transitions, and background, which is created from among the photos, videos, and other content stored on the user’s assistant-enabled device – in particular embodiments, the assistant system may understand the visual features of the memories, contextual information around them (e.g., meta information such as people in the scene, time, location, etc.), allowing for easy and intuitive navigation through connected memories – in addition, the users may manipulate the created story to their liking, .e.g., making it longer/shorter, reordering clips, adding background music/text, and visual effects; col. 46, lines 38-43 - the embodiments disclosed herein present a novel conversational tool to interactively create and edit montages from a personal media collection – in particular embodiments, the first user request may comprise a first utterance from the first user and the second user request may comprise a second utterance from the first user; col. 46, line 51- col. 47, line 12 – the embodiments disclosed herein collect C3, a TOD dialog dataset aimed at providing an intuitive conversational interface in which users may search through their media, create a video story with highlights, and edit clips hand-free, using natural language – Fig. 5 illustrates an example dialog 50 for conversational content creation – each dialog turn may be fully annotated with dialog acts and multimodal coreference labels, accompanied with its corresponding story montage snapshot; col. 51, lines 28-33 – the assistant system 140 may construct a user memory graph where the metadata from the user’s photo and video collections may be compiled together – the assistant system 140 may form natural connections among the photos and videos based on the metadata and the predicted visual features).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had obtained a multi-modal contextual input from a user input and identified events in, within each of a plurality of video frames of the video, based on analysis of the plurality of video frames and the multi-modal contextual input as disclosed by Kottur et al. in the method disclosed by Lim et al. in order to allow a user to create a video story with highlights and edit clips hands-free, using natural language. However, Lim et al. in view of Kottur et al. still fails to disclose identifying a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event; and obtaining a semantic relationship between identified primary event and the one or more secondary events.
Referring to the Park reference, Park discloses an electronic device for displaying a video, the electronic device comprising one or more processors configured to: identifying a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event (paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data); and obtaining a semantic relationship between identified primary event and the one or more secondary events (paragraph [0023] – the streaming service apparatus 200 divides the execution screen of the content into the main area and the auxiliary area according to the contents of the contents – the main area and the auxiliary area are set in consideration of the intention of the content creator; paragraph [0024] - the main domain corresponds to an essential part of the contents of the contents to grasp the contents of the contents - for example, if the content is a movie or a TV drama, it may be a person who is talking about the present dialogue, or a person or an object that is focused - in the case where the content is a streaming game, the user's character, the object selected by the user, and the target for proceeding the game can be the main areas; paragraph [0025] - the subarea corresponds to a portion of content that is less important than the main region of the content - for example, if the content is a movie or a TV drama, it may be an auxiliary person, simple background not related to the drama flow; paragraph [0033] – when the content is executed, the streaming device 200 divides the generated execution screen by executing the content (S204) – at this time, the streaming service apparatus 200 can divide the execution screen based on the contents of the contents. The location, size, and shape of the image are also set based on the contents of the content – the execution screen is divided into a main area and a sub area according to the contents of the contents – the main area and the auxiliary area of the content can be arbitrarily set by the content creator or the administrator of the streamlining service device 200 - in the case where the content is a movie or a TV drama, it is also possible to recognize a face of the actor or a close-up object by using a screen recognition technology and set the corresponding part as a main area – when the content is a game, the main area and the auxiliary area can be divided by referring to the game data).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had identified a primary event in a primary region and one or more secondary events in one or more secondary regions, within each of video frames of a video, based on an analysis of the video frames, wherein the one or more secondary events are contextually related to the primary event; and obtained a semantic relationship between identified primary event and the one or more secondary events as disclosed by Park in the device disclosed by Lim et al. in view of Kottur et al. in order to ensure that the regions of interest are further enhanced. Once Lim et al., Kottur et al., and Park are combined, the obtained frames matching the determined first aspect ratio disclosed by Lim et al. would include the primary event and the one or more secondary events located at the predicted positions in the obtained frames.
Claims 11 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Lim et al. in view of Kottur et al. in view of Park as applied to claims 1 and 12 above, and further in view of Yoon et al. (U.S. Patent Application Publication 2020/0027226).
Regarding claim 11, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claims 1 and 10, but fails to disclose that wherein the plurality of aesthetic effects comprises at least one of a depth effect, a pose change effect, a luminance effect, a lighting effect, or an audio effect associated with the primary region.
Referring to the Yoon et al. reference, Yoon et al. discloses a method for displaying a video, the method comprising: wherein the plurality of aesthetic effects comprises at least one of a depth effect, a pose change effect, a luminance effect, a lighting effect, or an audio effect associated with the primary region (paragraph [0077] – the applying module 230 may focus an object within a predetermined area from where the touch is sensed while applying an image effect (e.g., blurring) to at least one object other than the object in the previewed image – the applying module 230 may apply the image effect (e.g., blurring) to the object using the respective depth information corresponding to the at least one object – the image effect may include adjusting one of blur, color, brightness, mosaic, and resolution; claims 1-3 and 6).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had the plurality of aesthetic effects comprise at least one of a depth effect, a pose change effect, a luminance effect, a lighting effect, or an audio effect associated with the primary region as disclosed by Yoon et al. in the method disclosed by Lim et al. in view of Kottur et al. in view of Park in order to improve the overall look of the image.
Regarding claim 19, Lim et al. in view of Kottur et al. in view of Park discloses all of the limitations as previously discussed with respect to claims 12 and 18, but fails to disclose that wherein the plurality of aesthetic effects comprises at least one of a depth effect, a pose change effect, a luminance effect, a lighting effect, or an audio effect associated with the primary region.
Referring to the Yoon et al. reference, Yoon et al. discloses an electronic device for displaying a video, the electronic device comprising one or more processors configured to: wherein the plurality of aesthetic effects comprises at least one of a depth effect, a pose change effect, a luminance effect, a lighting effect, or an audio effect associated with the primary region (paragraph [0077] – the applying module 230 may focus an object within a predetermined area from where the touch is sensed while applying an image effect (e.g., blurring) to at least one object other than the object in the previewed image – the applying module 230 may apply the image effect (e.g., blurring) to the object using the respective depth information corresponding to the at least one object – the image effect may include adjusting one of blur, color, brightness, mosaic, and resolution; claims 1-3 and 6).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had the plurality of aesthetic effects comprise at least one of a depth effect, a pose change effect, a luminance effect, a lighting effect, or an audio effect associated with the primary region as disclosed by Yoon et al. in the device disclosed by Lim et al. in view of Kottur et al. in view of Park in order to improve the overall look of the image.
Allowable Subject Matter
Claims 3-7, 14, and 15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: Prior art, either alone or in combination, fails to teach or fairly suggest in combination with all of the other elements claimed:
The aspect ratio based method further comprises: obtaining the plurality of video frames and the plurality of multi-modal contextual inputs from at least one of the user input or the one or more applications in the electronic device; and performing the analysis of the plurality of video frames, wherein the analysis of the plurality of video frames comprises: determining a depth map based on RedGreenBlue-Depth (RGBD) data or RedGreenBlue (RGB) data in the obtained plurality of video frames; identifying key corners for each of the plurality of video frames based on the RGBD data or the RGB data; estimating a depth-aware optical flow comprising one or more flow points respective of each of the plurality of video frames, based on the key corners and the depth map; classifying similar depth-aware optical flows, using curve matching techniques, into one or more categories, wherein the one or more categories respectively correspond to one or more flow clusters; determining a first category among the one or more categories having a highest cardinality, wherein the highest cardinality corresponds to a highest number of optical flows in a cluster among one or more clusters; obtaining one or more convex hull points, encompassing the one or more flow points in each of the one or more clusters and the first category; and determining one or more bounding boxes enclosing each of the obtained one or more convex hull points, wherein the one or more bounding boxes comprise the primary region and the one or more secondary regions (dependent claim 3, which depends from claims 1 and 2; claims 4-7 depend from claim 3).
wherein the instructions, when executed by the one or more processors individually or collectively, cause the electronic device to: obtain the plurality of video frames and the plurality of multi-modal contextual inputs from at least one of the user input or the one or more applications in the electronic device; and determine a depth map based on RedGreenBlue-Depth (RGBD) data or RedGreenBlue (RGB) data in the obtained plurality of video frames; identify key corners for each of the plurality of video frames based on the RGBD data or the RGB data; estimate a depth-aware optical flow including one or more flow points respective of each of the plurality of video frames based on the key corners and the depth map; classify similar depth-aware optical flows, using curve matching techniques, into one or more categories, wherein the one or more categories corresponds to one or more flow clusters; determine a first category among the one or more categories having a highest cardinality, wherein the highest cardinality corresponds to a highest number of optical flows in a cluster among one or more clusters; obtain one or more convex hull points encompassing the one or more flow points in each of the one or more clusters and the first category; and determine one or more bounding boxes enclosing each of the obtained one or more convex hull points, wherein the one or more bounding boxes comprise the primary region and the one or more secondary regions (dependent claim 14, which depends from claims 12 and 13; claim 15 depends from claim 14).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HEATHER R JONES whose telephone number is (571)272-7368. The examiner can normally be reached Mon. - Fri.: 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William Vaughn can be reached at (571)272-3922. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HEATHER R JONES/Primary Examiner, Art Unit 2481
September 18, 2026