Prosecution Insights
Last updated: July 29, 2026
Application No. 18/640,694

IMAGE COMPRESSION APPARATUS AND METHOD

Non-Final OA §103
Filed
Apr 19, 2024
Priority
Oct 20, 2021 — RE 10-2021-0140544 +2 more
Examiner
ALFONSO, DENISE G
Art Unit
2662
Tech Center
2600 — Communications
Assignee
Hanwha Corporation
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
86 granted / 116 resolved
+12.1% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
15 currently pending
Career history
141
Total Applications
across all art units

Statute-Specific Performance

§101
0.6%
-39.4% vs TC avg
§103
90.8%
+50.8% vs TC avg
§102
6.6%
-33.4% vs TC avg
§112
1.2%
-38.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 116 resolved cases

Office Action

§103
DETAILED ACTIONS Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim this application being a Continuation of the International Application No. PCT/KR22/15993, filed on October 20, 2022, and benefit of foreign priority from Korean Patent Application No. KR10-2021-0140544 filed on October 20, 2021. Information Disclosure Statement The information disclosure statement (“IDS”) filed on 04/19/2024 was reviewed and the listed references were noted. Drawings The 10-page drawings have been considered and placed on record in the file. Status of Claims Claims 1-20 are pending. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-3 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kalva et al., (US 2010/0158099 A1, published 06/24/2010), hereinafter referred to as Kalva, in view of Pan et al., (US 2017/0134754 A1, published 05/11/2017), hereinafter referred to as Pan. Claim 1 Kalva discloses an image compression method (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”) performed by an apparatus (Kalva, Fig. 1B) comprising at least one processor and at least one memory (Kalva, [0041], “The apparatus 108 may be implemented on and/or in conjunction with a computing device 107 which, as discussed above, may comprise a processor, memory, computer-readable media, communications interfaces, a human-machine interface (HMI) 177, and the like.”) that stores instructions executable by the at least one processor (Kalva, [0041], “The apparatus 108 may be implemented on and/or in conjunction with a computing device 107 which, as discussed above, may comprise a processor, memory, computer-readable media, communications interfaces, a human-machine interface (HMI) 177, and the like..”), the method comprising: receiving an event information of a captured image (Kalva, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like. In the FIG. 1A example, the ROI identified by the content-aware metadata 131 identify the time a particular player enters the segment 103”); encoding an image frame from the captured image (Kalva, [0037], “the Content Adaptive Video Encoder 157 may provide one or more alternative encoding(s), which may comprise alternative encodings of the entire video content 140 and/or portions thereof (e.g., particular regions, segments, or the like)”); generating a meta-frame by encoding a mapping (Kalva, [0040], “The separate content-aware metadata may flow to an indexing and/or search component for use in indexing and/or classifying the encoded video content 146”, [0019], “Content-aware metadata may be embodied as text data (e.g., UTF-8 encoded strings), formatted data (e.g., XML), compressed binary data, or the like.”, the mapping table example in Fig. 2 of the Specification of the application shows the object type and the binary representation of the object, Kalva discloses encoding the metadata into a compressed binary and indexing them which is analogous to the mapping) corresponding to the event information (Kalva, [0039], “The Content-Aware Metadata Encoder 159 may be configured to embed the content-aware metadata in the encoded video content 146 (e.g., embed the content-aware metadata in the video bitstream). The embedding may provide for a tight coupling of the encoded video content and the content-aware metadata that is independent of the mechanism used to playback and/or transport the encoded content 146”, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content), regions having a particular color, text displayed within or otherwise obtained from the video (e.g., as a graphical element of the video, as sub-title information, menu items, text obtained from or by processing an audio track, or the like), shape, edge or other identifying characteristic, scene change information, scene fingerprint, scene complexity, identification of background and foreground regions, texture descriptors, and the like”, [0030], “or example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”); generating a transmission packet by combining the meta-frame with the encoded image frame (Kalva, [0077], “The encoded video content produced by the video stream parser 480 and/or Content Adaptive Encoder 455 may flow to the Multiplexer and Packetizer module 482, which may combine encoded video content with IVMA data authored by the author 105 (in textual or binary, compressed format) into a single stream and/or as a set of distributed IVMA data instances (as defined in the schema 420). The IVMA data authored by the author 105 (using the IVMA Authoring module 415) may be processed by the IVMA Data Parser module 430 and/or IVMA Data Encoder modules 435. As discussed above, the IVMA Data Parser module 430 may be configured to segment the IVMA data into chunks. Each chunk of IVMA data may comprise one or more IVMA data elements or instances (e.g., in accordance with the schema 420 described below). The chunks may be binary encoded and/or compressed by the IVMA Data Encoder module 435. Alternatively, the IVMA data may bypass the IVMA Data Parser module 430 and/or IVMA Data Encoder module 435, and flow directly to the Multiplexer and Packetizer module 482 and/or made available from the apparatus 400 as text..”, as shown in Fig. 1B, both the encoded video content and encoded metadata is outputted by the Content Adaptive Encoder); and transmitting the generated transmission packet (Kalva, [0078], “The Multiplexer and Packetizer module 482 may be further configured to packetize the multiplexed stream for transmission on a network. The packet and/or transmission frame size may be selected according to the performance and/or capabilities of the network infrastructure used to transmit the data.”), wherein the mapping table (Kalva, [0040], “The separate content-aware metadata may flow to an indexing and/or search component for use in indexing and/or classifying the encoded video content 146”, [0019], “Content-aware metadata may be embodied as text data (e.g., UTF-8 encoded strings), formatted data (e.g., XML), compressed binary data, or the like.”, the mapping table example in Fig. 2 of the Specification of the application shows the object type and the binary representation of the object, Kalva discloses encoding the metadata into a compressed binary and indexing them which is analogous to the mapping table) comprises a first mapping (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content),” [0035], “For example, the machine vision techniques may be used to classify a shape identified by an image compression technique (e.g., identify the shape as a racecar and/or as a particular racecar, as a baseball player, and so on)”) and a second mapping (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content)”, under BRI, the motion of the object is considered a situation of the object, [0030], “the frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”). Kalva discloses mapping a binary representation for each metadata including the classification of the type of the object or the motion or action of the ROI in the frame. Kalva does not explicitly disclose generating a meta-frame by encoding a mapping table corresponding to the event information, wherein the mapping table comprises a first mapping table for encoding an object type for classifying at least one object included in the event information and a second mapping table for encoding a situation class for classifying a situation of the at least one object. However, Pan teaches generating a meta-frame by encoding a mapping table corresponding to the event information (Pan, Fig. 2, [0036], “FIG. 2, shows examples of video data representations for an object and background portions of a video image frame according to an aspect of the present disclosure.”, [0035], “additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, [0031], “Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions”), wherein the mapping table comprises a first mapping table for encoding an object type for classifying at least one object included in the event information (Fig. 2, table 210 is analogous to the mapping table, [0035], “ The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier”, [0031], “Objects that are detected by the object detection algorithm are considered to be foreground objects. Portions of the image that are not in a detected object are considered to be background portions of the image. Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions.”) and a second mapping table for encoding a situation class for classifying a situation of the at least one object (Fig. 2, table 210, [0035], “According to an aspect of the present disclosure, additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, motion information. Kalva and Pan are both considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method as taught by Kalva to incorporate the teachings of Pan of mapping a binary representation for each metadata including the classification of the type of the object or the motion or action of the ROI in the frame. Kalva does not explicitly disclose generating a meta-frame by encoding a mapping table corresponding to the event information, wherein the mapping table comprises a first mapping table for encoding an object type for classifying at least one object included in the event information and a second mapping table for encoding a situation class for classifying a situation of the at least one object. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been for faster access to information (Pan, [0007]). Claim 2 The combination of Kalva in view of Pan discloses the image compression method of claim 1 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”) wherein the object type in the first mapping table (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content),”) has a first priority and a simpler code is mapped to an object type with a higher first priority (Pan, [0031], “Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions.”, Kalva teaches encoding the ROI objects in the image, and Pan teaches that the ROI or foreground objects are given higher priority), and the situation class in the second mapping table has a second priority and a simpler code is mapped to a situation class with a higher second priority (Pan, [0033], “Each detected object and background portion are encoded as a wavelet pyramid. Some sub-bands of the background portions of an image have low priority, thus a low bit rate can be used to encode the low priority sub-bands representing the background portions.”, [0035], “additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, Pan teaches that the additional information or background portions of metadata which includes the motion information has less priority than the foreground object, motion information is analogous to the situation of the object). The proposed combination as well as the motivation for combining the Kalva and Pan references presented in the rejection of Claim 1, apply to Claim 2 and are incorporated herein by reference. Thus, the method recited in Claim 2 is met by Kalva and Pan. Claim 3 The combination of Kalva in view of Pan discloses the image compression method of claim 3 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the meta-frame comprises a field in which the first mapping table is recorded (Pan, Fig. 2, table 206), a field in which the second mapping table is recorded, (Pan, Fig. 2, table 206), a field in which a probability that the object type (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content),” [0035], “For example, the machine vision techniques may be used to classify a shape identified by an image compression technique (e.g., identify the shape as a racecar and/or as a particular racecar, as a baseball player, and so on)”) of the first mapping table is correct is recorded (Pan, [0040], “Furthermore, other features, such as associated track ID, lifetime, and/or target confidence level, for example, can also be encoded and associated with the encoding of each corresponding object to facilitate faster video content query and decision making.”), and a field in which a probability in which a situation class (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content)”, under BRI, the motion of the object is considered a situation of the object, [0030], “the frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”) of the second mapping table is correct is recorded (Pan, [0040], “Furthermore, other features, such as associated track ID, lifetime, and/or target confidence level, for example, can also be encoded and associated with the encoding of each corresponding object to facilitate faster video content query and decision making.”). The proposed combination as well as the motivation for combining the Kalva and Pan references presented in the rejection of Claim 1, apply to Claim 2 and are incorporated herein by reference. Thus, the method recited in Claim 2 is met by Kalva and Pan. Claim 18 Kalva discloses an image compression apparatus (Kalva, Fig. 1B) comprising: at least one memory configured to store thereon one or more computer program codes (Kalva, [0041], “The apparatus 108 may be implemented on and/or in conjunction with a computing device 107 which, as discussed above, may comprise a processor, memory, computer-readable media, communications interfaces, a human-machine interface (HMI) 177, and the like.”); and at least one processor configures to access the at least one memory and operate according to the one or more computer program codes, wherein the one or more computer program codes are configured to cause at least one processor to perform (Kalva, [0041], “The apparatus 108 may be implemented on and/or in conjunction with a computing device 107 which, as discussed above, may comprise a processor, memory, computer-readable media, communications interfaces, a human-machine interface (HMI) 177, and the like.”): acquiring an event information of a captured image (Kalva, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like. In the FIG. 1A example, the ROI identified by the content-aware metadata 131 identify the time a particular player enters the segment 103”); encoding an image frame from the captured image (Kalva, [0037], “the Content Adaptive Video Encoder 157 may provide one or more alternative encoding(s), which may comprise alternative encodings of the entire video content 140 and/or portions thereof (e.g., particular regions, segments, or the like)”); generating a meta-frame by encoding a mapping (Kalva, [0040], “The separate content-aware metadata may flow to an indexing and/or search component for use in indexing and/or classifying the encoded video content 146”, [0019], “Content-aware metadata may be embodied as text data (e.g., UTF-8 encoded strings), formatted data (e.g., XML), compressed binary data, or the like.”, the mapping table example in Fig. 2 of the Specification of the application shows the object type and the binary representation of the object, Kalva discloses encoding the metadata into a compressed binary and indexing them which is analogous to the mapping) corresponding to the event information (Kalva, [0039], “The Content-Aware Metadata Encoder 159 may be configured to embed the content-aware metadata in the encoded video content 146 (e.g., embed the content-aware metadata in the video bitstream). The embedding may provide for a tight coupling of the encoded video content and the content-aware metadata that is independent of the mechanism used to playback and/or transport the encoded content 146”, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content), regions having a particular color, text displayed within or otherwise obtained from the video (e.g., as a graphical element of the video, as sub-title information, menu items, text obtained from or by processing an audio track, or the like), shape, edge or other identifying characteristic, scene change information, scene fingerprint, scene complexity, identification of background and foreground regions, texture descriptors, and the like”, [0030], “or example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”); generating a transmission packet by combining the meta-frame with the encoded image frame (Kalva, [0077], “The encoded video content produced by the video stream parser 480 and/or Content Adaptive Encoder 455 may flow to the Multiplexer and Packetizer module 482, which may combine encoded video content with IVMA data authored by the author 105 (in textual or binary, compressed format) into a single stream and/or as a set of distributed IVMA data instances (as defined in the schema 420). The IVMA data authored by the author 105 (using the IVMA Authoring module 415) may be processed by the IVMA Data Parser module 430 and/or IVMA Data Encoder modules 435. As discussed above, the IVMA Data Parser module 430 may be configured to segment the IVMA data into chunks. Each chunk of IVMA data may comprise one or more IVMA data elements or instances (e.g., in accordance with the schema 420 described below). The chunks may be binary encoded and/or compressed by the IVMA Data Encoder module 435. Alternatively, the IVMA data may bypass the IVMA Data Parser module 430 and/or IVMA Data Encoder module 435, and flow directly to the Multiplexer and Packetizer module 482 and/or made available from the apparatus 400 as text..”, as shown in Fig. 1B, both the encoded video content and encoded metadata is outputted by the Content Adaptive Encoder); and transmitting the generated transmission packet (Kalva, [0078], “The Multiplexer and Packetizer module 482 may be further configured to packetize the multiplexed stream for transmission on a network. The packet and/or transmission frame size may be selected according to the performance and/or capabilities of the network infrastructure used to transmit the data.”), wherein the mapping table (Kalva, [0040], “The separate content-aware metadata may flow to an indexing and/or search component for use in indexing and/or classifying the encoded video content 146”, [0019], “Content-aware metadata may be embodied as text data (e.g., UTF-8 encoded strings), formatted data (e.g., XML), compressed binary data, or the like.”, the mapping table example in Fig. 2 of the Specification of the application shows the object type and the binary representation of the object, Kalva discloses encoding the metadata into a compressed binary and indexing them which is analogous to the mapping table) comprises a first mapping (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content),” [0035], “For example, the machine vision techniques may be used to classify a shape identified by an image compression technique (e.g., identify the shape as a racecar and/or as a particular racecar, as a baseball player, and so on)”) and a second mapping (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content)”, under BRI, the motion of the object is considered a situation of the object, [0030], “the frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”). Kalva discloses mapping a binary representation for each metadata including the classification of the type of the object or the motion or action of the ROI in the frame. Kalva does not explicitly disclose generating a meta-frame by encoding a mapping table corresponding to the event information, wherein the mapping table comprises a first mapping table for encoding an object type for classifying at least one object included in the event information and a second mapping table for encoding a situation class for classifying a situation of the at least one object. However, Pan teaches generating a meta-frame by encoding a mapping table corresponding to the event information (Pan, Fig. 2, [0036], “FIG. 2, shows examples of video data representations for an object and background portions of a video image frame according to an aspect of the present disclosure.”, [0035], “additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, [0031], “Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions”), wherein the mapping table comprises a first mapping table for encoding an object type for classifying at least one object included in the event information (Fig. 2, table 210 is analogous to the mapping table, [0035], “ The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier”, [0031], “Objects that are detected by the object detection algorithm are considered to be foreground objects. Portions of the image that are not in a detected object are considered to be background portions of the image. Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions.”) and a second mapping table for encoding a situation class for classifying a situation of the at least one object (Fig. 2, table 210, [0035], “According to an aspect of the present disclosure, additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, motion information. Kalva and Pan are both considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus as taught by Kalva to incorporate the teachings of Pan of mapping a binary representation for each metadata including the classification of the type of the object or the motion or action of the ROI in the frame. Kalva does not explicitly disclose generating a meta-frame by encoding a mapping table corresponding to the event information, wherein the mapping table comprises a first mapping table for encoding an object type for classifying at least one object included in the event information and a second mapping table for encoding a situation class for classifying a situation of the at least one object. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been for faster access to information (Pan, [0007]). Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Kalva in view of Pan in further view of Shahbazi et al., (US 2022/0383881 A1, filed 05/27/2021), hereinafter referred to as Shahbazi. Claim 4 The combination of Kalva in view of Pan discloses the image compression method of claim 3 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the meta-frame is generated only for at least one image frame having the event information, among image frames (Kalva, [0021], “The tight-coupling described above may allow content-aware metadata to be associated with the video content at the frame, scene, and/or video object level”). The combination of Kalva in view of Pan does not explicitly disclose whether an image frame has a meta-frame is indicated by a flag bit. However, Shahbazi teaches whether an image frame has a meta-frame is indicated by a flag bit (Shahbazi, [0114], “The metadata flags field 705 includes one or more metadata flags that indicate the presence of various types of metadata 722 in the media payload 720. For example, a first flag 709 may indicate the presence of motion data in the metadata 722, such as the motion data from the one or more motion sensors 630 that may be sent along with the encoded audio data to the second device 604 of FIG. 6. As another example, a second flag 710 may indicate the presence of link data in the metadata 722, such as the link data 168. One or more additional metadata flags may be included in the metadata flags field 705, such as a third flag to indicate that spatial information (e.g., based on beamforming associated with speech capture at the second device 604) is included in the metadata 722. In some implementations, each of the flags in the metadata flags field 705 is represented as a single bit, and a recipient of the media packet 700 can determine the length of the metadata 722 in the media payload 720 based on the particular bits in the metadata flags field 705 that are set.”). Kalva, Pan, and Shahbazi are all considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method as taught by Kalva and Pan to incorporate the teachings of Shahbazi whether an image frame has a meta-frame is indicated by a flag bit. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have to help indicate various types of metadata (Shahbazi, [0114]). Claims 7-10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Kalva in view of Teuton et al., (US 2016/0314354 A1, published 10/27/2016), hereinafter referred to as Teuton. Claim 7 Kalva discloses an image compression method (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”) performed by an apparatus (Kalva, Fig. 1B) comprising at least one processor and at least one memory (Kalva, [0041], “The apparatus 108 may be implemented on and/or in conjunction with a computing device 107 which, as discussed above, may comprise a processor, memory, computer-readable media, communications interfaces, a human-machine interface (HMI) 177, and the like.”) that stores instructions executable by the at least one processor (Kalva, [0041], “The apparatus 108 may be implemented on and/or in conjunction with a computing device 107 which, as discussed above, may comprise a processor, memory, computer-readable media, communications interfaces, a human-machine interface (HMI) 177, and the like..”), the method comprising: generating an event information of a captured image (Kalva, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like. In the FIG. 1A example, the ROI identified by the content-aware metadata 131 identify the time a particular player enters the segment 103”); encoding an image frame from the captured image (Kalva, [0037], “the Content Adaptive Video Encoder 157 may provide one or more alternative encoding(s), which may comprise alternative encodings of the entire video content 140 and/or portions thereof (e.g., particular regions, segments, or the like)”); generating a meta-frame by encoding a mapping - (Kalva, [0040], “The separate content-aware metadata may flow to an indexing and/or search component for use in indexing and/or classifying the encoded video content 146”, [0019], “Content-aware metadata may be embodied as text data (e.g., UTF-8 encoded strings), formatted data (e.g., XML), compressed binary data, or the like.”, the mapping table example in Fig. 2 of the Specification of the application shows the object type and the binary representation of the object, Kalva discloses encoding the metadata into a compressed binary and indexing them which is analogous to the mapping) corresponding to the event information (Kalva, [0039], “The Content-Aware Metadata Encoder 159 may be configured to embed the content-aware metadata in the encoded video content 146 (e.g., embed the content-aware metadata in the video bitstream). The embedding may provide for a tight coupling of the encoded video content and the content-aware metadata that is independent of the mechanism used to playback and/or transport the encoded content 146”, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content), regions having a particular color, text displayed within or otherwise obtained from the video (e.g., as a graphical element of the video, as sub-title information, menu items, text obtained from or by processing an audio track, or the like), shape, edge or other identifying characteristic, scene change information, scene fingerprint, scene complexity, identification of background and foreground regions, texture descriptors, and the like”, [0030], “or example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”); generating a meta-frame by (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”) and artificial intelligence (AI) data (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”) , the MD data and the AI data corresponding to the generated event information (Kalva, [0034], “The Content Adaptive Preprocessor 150 may analyze the video content 140 using video compression and/or machine vision techniques, each of which may result in the identification of content-aware metadata. The video compression information may include scene change information, scene complexity information, shape identification, local motion vectors, texture descriptors, region of interest, and the like.”, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like”); generating a transmission packet by combining the meta-frame with the encoded image frame (Kalva, [0077], “The encoded video content produced by the video stream parser 480 and/or Content Adaptive Encoder 455 may flow to the Multiplexer and Packetizer module 482, which may combine encoded video content with IVMA data authored by the author 105 (in textual or binary, compressed format) into a single stream and/or as a set of distributed IVMA data instances (as defined in the schema 420). The IVMA data authored by the author 105 (using the IVMA Authoring module 415) may be processed by the IVMA Data Parser module 430 and/or IVMA Data Encoder modules 435. As discussed above, the IVMA Data Parser module 430 may be configured to segment the IVMA data into chunks. Each chunk of IVMA data may comprise one or more IVMA data elements or instances (e.g., in accordance with the schema 420 described below). The chunks may be binary encoded and/or compressed by the IVMA Data Encoder module 435. Alternatively, the IVMA data may bypass the IVMA Data Parser module 430 and/or IVMA Data Encoder module 435, and flow directly to the Multiplexer and Packetizer module 482 and/or made available from the apparatus 400 as text..”, as shown in Fig. 1B, both the encoded video content and encoded metadata is outputted by the Content Adaptive Encoder); and transmitting the generated transmission packet (Kalva, [0078], “The Multiplexer and Packetizer module 482 may be further configured to packetize the multiplexed stream for transmission on a network. The packet and/or transmission frame size may be selected according to the performance and/or capabilities of the network infrastructure used to transmit the data.”), wherein the MD data is a low-level event information obtained through motion detection between a plurality of image frames (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like”, motion vector estimation is a lower level data compared to the output of a machine learning technique), and the AI data is a high-level event information obtained through AI learning (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”, the output of machine learning is considered more high-level than a motion vector estimation). Kalva does not explicitly disclose generating a meta-frame by losslessly encoding motion detection (MD) data and artificial intelligence (AI) data. However, Teuton teaches generating a meta-frame (Teuton, Abstract, “, a task-specific processing component, such as a video encoder, is used to generate task-specific metadata. When the data set includes video frames, metadata can include information regarding motion of image elements between frames, or other differences between frames. A feature of the data set, such as an event, can be identified or classified based on the metadata.”) by losslessly encoding (Teuton, [0061], “consider a feature detection/classification component 105 in which the task-specific processing component 125 is a video encoder configured to generate encoded (compressed) video data using a compression or encoding process (which can be lossless or lossy)”) motion detection (MD) data (Teuton, [0135], “Metadata generated by the video encoder 440 can also include motion vector data that describes the movement of the macroblocks, including the magnitude and direction of motion between frames 510”) and artificial intelligence (AI) data (Teuton, [0083], “the metadata analysis component 155 can be, or can include, a machine learning software component.”). Kalva and Teuton are all considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method as taught by Kalva to incorporate the teachings of Teuton of generating a meta-frame by losslessly encoding motion detection (MD) data and artificial intelligence (AI) data. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have to exploit redundancies in video data (Teuton, [0045]). Claim 8 The combination of Kalva in view of Teuton discloses the image compression method of claim 7 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the generating the meta-frame comprises generating the meta-frame by selectively losslessly encoding (Kalva, [0033], “the apparatus 104 may be configured to identify and embed content-aware metadata into encoded video content 146”, Teuton, [0061], “consider a feature detection/classification component 105 in which the task-specific processing component 125 is a video encoder configured to generate encoded (compressed) video data using a compression or encoding process (which can be lossless or lossy)”) at least one of the low-level event information or the high-level event information (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like”) at a request of an image restoration device (Kalva, [0070], “The content-aware metadata may be provided to the IVM Application Interpreter and Asset Request Generator module 362, which, as discussed above, may use the content-aware metadata to update the IVM application (e.g., by a direct communication path therebetween (not shown) and/or through a communication path provided by the IVM Application Components Renderer and Compositor module 365). As discussed above, the content-aware metadata may be used by the IVM Application Interpreter and Asset Request Generator module 362 to determine rendering, composition, and/or user-interactivity features of the IVM application as defined by the IVMA data.”). Claim 9 The combination of Kalva in view of Teuton discloses the image compression method of claim 7 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the MD data comprises a first data field for identifying an image frame comprising an area in which a motion was detected (Kalva, [0030], “[0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like”) , a second data field for recording a time when the motion was detected (Kalva, [0029], “s used in this context, the ROI metadata 131 may specify a time segment of the video content and/or a particular set of video frames as opposed to a region within a video display area.”), and a third data field for recording a location of the area in which the motion was detected in the image frame (Kalva, [0023]”A region layer may be assigned various properties, including, but not limited to: region color, texture, motion, location (within one or more frames), shape, and the like. Each region may be assigned an identifier. The region layers (as well as their associated properties) may be made available by the encoder as content-aware metadata.”). Claim 10 The combination of Kalva in view of Teuton discloses the image compression method of claim 9 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the MD data further comprises a fourth data field for recording at least one of a horizontal size or a vertical size of the area in which the motion was detected (Kalva, [0059], “the IVA application may use the content-aware metadata to cause the display to be placed in the vicinity of the racecar object and/or may cause the display to follow the movement of the racecar as it moves around the track (e.g., using content-aware metadata identifying the position and/or local motion vectors of the object)”, the vector has the distance information between points which can show the horizontal or vertical distance or size of the motion). Claim 17 The combination of Kalva in view of Teuton discloses the image compression method of claim 7 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the losslessly encoding (Kalva, [0033], “the apparatus 104 may be configured to identify and embed content-aware metadata into encoded video content 146”, Teuton, [0061], “consider a feature detection/classification component 105 in which the task-specific processing component 125 is a video encoder configured to generate encoded (compressed) video data using a compression or encoding process (which can be lossless or lossy)”) the MD data and the AI data (Kalva, [0034], “The Content Adaptive Preprocessor 150 may analyze the video content 140 using video compression and/or machine vision techniques, each of which may result in the identification of content-aware metadata. The video compression information may include scene change information, scene complexity information, shape identification, local motion vectors, texture descriptors, region of interest, and the like.”, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like”); is performed by an entropy coding unit in a video encoder which encodes the image frame (Teuton, [0061], “consider a feature detection/classification component 105 in which the task-specific processing component 125 is a video encoder configured to generate encoded (compressed) video data using a compression or encoding process (which can be lossless or lossy)”, [0167], “Motion vectors are calculated during a phase of the encoding process called motion estimation. A video encoder can calculate differences, or residual values, between the original sample values of a macroblock and the motion-compensated predicted values for that macroblock. The residual values can be further compressed using a DCT, quantization, and entropy coding.”). The proposed combination as well as the motivation for combining the Kalva and Teuton references presented in the rejection of Claim 7, apply to Claim 17 and are incorporated herein by reference. Thus, the method recited in Claim 17 is met by Kalva and Teuton. Claims 11-13 are rejected under 35 U.S.C. 103 as being unpatentable over Kalva in view of Teuton in further view of Pan. Claim 11 The combination of Kalva in view of Teuton discloses the image compression method of claim 7 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the AI data (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like) comprises a first comprises a first mapping (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content),” [0035], “For example, the machine vision techniques may be used to classify a shape identified by an image compression technique (e.g., identify the shape as a racecar and/or as a particular racecar, as a baseball player, and so on)”) and a second mapping (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content)”, under BRI, the motion of the object is considered a situation of the object, [0030], “the frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”). Kalva discloses mapping a binary representation for each metadata including the classification of the type of the object or the motion or action of the ROI in the frame. Kalva does not explicitly disclose generating a meta-frame by encoding a mapping table corresponding to the event information, wherein the AI data comprises a first mapping table for encoding an object type for classifying at least one object included in the event information and a second mapping table for encoding a situation class for classifying a situation of the at least one object. However, Pan teaches generating a meta-frame by encoding a mapping table corresponding to the event information (Pan, Fig. 2, [0036], “FIG. 2, shows examples of video data representations for an object and background portions of a video image frame according to an aspect of the present disclosure.”, [0035], “additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, [0031], “Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions”), wherein the AI data (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like) comprises a first mapping table for encoding an object type for classifying at least one object included in the event information (Pan, Fig. 2, table 210 is analogous to the mapping table, [0035], “ The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier”, [0031], “Objects that are detected by the object detection algorithm are considered to be foreground objects. Portions of the image that are not in a detected object are considered to be background portions of the image. Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions.”) and a second mapping table for encoding a situation class for classifying a situation of the at least one object (Pan, Fig. 2, table 210, [0035], “According to an aspect of the present disclosure, additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, motion information. Kalva, Teuton, and Pan are all considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method as taught by Kalva to incorporate the teachings of Pan of mapping a binary representation for each metadata including the classification of the type of the object or the motion or action of the ROI in the frame. Kalva does not explicitly disclose generating a meta-frame by encoding a mapping table corresponding to the event information, wherein the AI data comprises a first mapping table for encoding an object type for classifying at least one object included in the event information and a second mapping table for encoding a situation class for classifying a situation of the at least one object. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been for faster access to information (Pan, [0007]). Claim 12 The combination of Kalva in view of Teuton in view of Pan discloses the image compression method of claim 11 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”) wherein the object type in the first mapping table (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content),”) has a first priority and a simpler code is mapped to an object type with a higher first priority (Pan, [0031], “Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions.”, Kalva teaches encoding the ROI objects in the image, and Pan teaches that the ROI or foreground objects are given higher priority), and the situation class in the second mapping table has a second priority and a simpler code is mapped to a situation class with a higher second priority (Pan, [0033], “Each detected object and background portion are encoded as a wavelet pyramid. Some sub-bands of the background portions of an image have low priority, thus a low bit rate can be used to encode the low priority sub-bands representing the background portions.”, [0035], “additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, Pan teaches that the additional information or background portions of metadata which includes the motion information has less priority than the foreground object, motion information is analogous to the situation of the object). The proposed combination as well as the motivation for combining the Kalva, Teuton, and Pan references presented in the rejection of Claim 11, apply to Claim 12 and are incorporated herein by reference. Thus, the method recited in Claim 12 is met by Kalva, Teuton, and Pan. Claim 13 The combination of Kalva in view of Teuton in view of Pan discloses the image compression method of claim 12 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”) wherein the meta-frame comprises a field in which the first mapping table is recorded (Pan, Fig. 2, table 206), a field in which the second mapping table is recorded, (Pan, Fig. 2, table 206), a field in which a probability that the object type (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content),” [0035], “For example, the machine vision techniques may be used to classify a shape identified by an image compression technique (e.g., identify the shape as a racecar and/or as a particular racecar, as a baseball player, and so on)”) of the first mapping table is correct is recorded (Pan, [0040], “Furthermore, other features, such as associated track ID, lifetime, and/or target confidence level, for example, can also be encoded and associated with the encoding of each corresponding object to facilitate faster video content query and decision making.”), and a field in which a probability in which a situation class (Kalva, [0019], “content-aware metadata may refer to information that describes or identifies a video content feature, including, but not limited to: a region of interest (ROI) within the video content (e.g., a particular portion or encoding aspect of a video display region, one or more video frames, etc.), an object within the content (e.g., shape, color region, etc.), motion characteristics of video objects (e.g., local motion vectors of a shape within the video content)”, under BRI, the motion of the object is considered a situation of the object, [0030], “the frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like.”) of the second mapping table is correct is recorded (Pan, [0040], “Furthermore, other features, such as associated track ID, lifetime, and/or target confidence level, for example, can also be encoded and associated with the encoding of each corresponding object to facilitate faster video content query and decision making.”). The proposed combination as well as the motivation for combining the Kalva, Teuton, and Pan references presented in the rejection of Claim 11, apply to Claim 13 and are incorporated herein by reference. Thus, the method recited in Claim 13 is met by Kalva, Teuton, and Pan. Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Kalva in view of Teuton in view of Shahbazi. Claim 14 The combination of Kalva in view of Teuton discloses the image compression method of claim 7 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the meta-frame is generated only for at least one image frame having the event information, among image frames (Kalva, [0021], “The tight-coupling described above may allow content-aware metadata to be associated with the video content at the frame, scene, and/or video object level”). The combination of Kalva in view of Teuton does not explicitly disclose whether an image frame has a meta-frame is indicated by a flag bit. However, Shahbazi teaches whether an image frame has a meta-frame is indicated by a flag bit (Shahbazi, [0114], “The metadata flags field 705 includes one or more metadata flags that indicate the presence of various types of metadata 722 in the media payload 720. For example, a first flag 709 may indicate the presence of motion data in the metadata 722, such as the motion data from the one or more motion sensors 630 that may be sent along with the encoded audio data to the second device 604 of FIG. 6. As another example, a second flag 710 may indicate the presence of link data in the metadata 722, such as the link data 168. One or more additional metadata flags may be included in the metadata flags field 705, such as a third flag to indicate that spatial information (e.g., based on beamforming associated with speech capture at the second device 604) is included in the metadata 722. In some implementations, each of the flags in the metadata flags field 705 is represented as a single bit, and a recipient of the media packet 700 can determine the length of the metadata 722 in the media payload 720 based on the particular bits in the metadata flags field 705 that are set.”). Kalva, Teuton and Shahbazi are all considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method as taught by Kalva to incorporate the teachings of Shahbazi whether an image frame has a meta-frame is indicated by a flag bit. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have to help indicate various types of metadata (Shahbazi, [0114]). Claims 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kalva in view of Pan in view of Teuton. Claim 19 The combination of Kalva in view of Pan discloses the image compression apparatus of claim 18 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the generating the meta-frame comprises generating the meta-frame by (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like), the AI data being a high-level event information obtained through AI learning (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”, the output of machine learning is considered more high-level than a motion vector estimation), and wherein the AI data comprises the first mapping table and the second mapping table (Pan, Fig. 2, [0036], “FIG. 2, shows examples of video data representations for an object and background portions of a video image frame according to an aspect of the present disclosure.”, [0035], “additional information that is descriptive of each compressed object and/or background portion is encoded and stored as metadata in association with the spatially compressed and/or temporally compressed representation of the corresponding object or background portion. The additional information may include features of an object or other attributes of the compressed video data associated with the object such as an object detection identifier, object detection method, detection value, track ID, motion information, and track confidence.”, [0031], “Foreground objects are generally given higher priority than the background portions of the image and are encoded using more encoding bits than the background portions”). The combination of Kalva in view of Pan does not explicitly disclose generating the meta-frame by losslessly encoding artificial intelligence (AI) data. However, Teuton teaches generating the meta-frame (Teuton, Abstract, “, a task-specific processing component, such as a video encoder, is used to generate task-specific metadata. When the data set includes video frames, metadata can include information regarding motion of image elements between frames, or other differences between frames. A feature of the data set, such as an event, can be identified or classified based on the metadata.”) by losslessly encoding (Teuton, [0061], “consider a feature detection/classification component 105 in which the task-specific processing component 125 is a video encoder configured to generate encoded (compressed) video data using a compression or encoding process (which can be lossless or lossy)”) artificial intelligence (AI) data (Teuton, [0083], “the metadata analysis component 155 can be, or can include, a machine learning software component.”). Kalva, Pan, and Teuton are all considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method as taught by Kalva to incorporate the teachings of Teuton of generating the meta-frame by losslessly encoding artificial intelligence (AI) data. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have to exploit redundancies in video data (Teuton, [0045]). Claim 20 The combination of Kalva in view of Pan discloses the image compression apparatus of claim 18 (Kalva, [0038], “The compressing and/or encoding implemented by the Content Adaptive Video Encoder 157 may yield additional content-aware metadata, such as global and local motion parameters, region/object shape actually encoded, texture descriptors, etc. In some embodiments, the Content Adaptive Video Encoder 157 may receive the metadata encoding conditions 144 to identify content-aware metadata for inclusion in the encoded content asset”), wherein the generating the meta-frame comprises generating the meta-frame by further (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”), the MD data being a low-level event information obtained through motion detection between a plurality of image frames (Kalva, [0028], “The ROI metadata 121 may be determined using an automated, machine-learning technique, such as pattern matching, template matching, optical character recognition (OCR), or the like. Alternatively, or in addition, video encoding and compression techniques may be used, such as shape identification, motion vector estimation, and the like”, [0030], “The frames specified by the ROI metadata 131 may correspond to a scene change (e.g., as identified by a compression, encoding, or machine vision operation), a scene fingerprint, or the like. The content-aware metadata 131 may identify a particular set of frames that may be of interest in video content. For example, the ROI identified by the metadata 131 may correspond to advertising insertion (e.g., during a lull in the action of a baseball game, such as the seventh inning stretch), to a particular event occurring in the content (e.g., a player hitting a home run), the entrance of a particular player onto the field, the presence of an object within the field of view (such as a billboard), an offensive scene, or the like”, motion vector estimation is a lower level data compared to the output of a machine learning technique). The combination of Kalva in view of Pan does not explicitly disclose generating a meta-frame by losslessly encoding motion detection (MD) data. However, Teuton teaches generating a meta-frame (Teuton, Abstract, “, a task-specific processing component, such as a video encoder, is used to generate task-specific metadata. When the data set includes video frames, metadata can include information regarding motion of image elements between frames, or other differences between frames. A feature of the data set, such as an event, can be identified or classified based on the metadata.”) by losslessly encoding (Teuton, [0061], “consider a feature detection/classification component 105 in which the task-specific processing component 125 is a video encoder configured to generate encoded (compressed) video data using a compression or encoding process (which can be lossless or lossy)”) motion detection (MD) data (Teuton, [0135], “Metadata generated by the video encoder 440 can also include motion vector data that describes the movement of the macroblocks, including the magnitude and direction of motion between frames 510”). Kalva, Pan, and Teuton are all considered to be analogous to the claimed invention because they are in the same field of metadata. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method as taught by Kalva to incorporate the teachings of Teuton of generating a meta-frame by losslessly encoding motion detection (MD) data. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have to exploit redundancies in video data (Teuton, [0045]). Allowable Subject Matter Claims 5 and 15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The claimed features such as “receiving first event information from a first event analysis source and receiving second event information from a second event analysis source, and wherein the generating the meta-frame comprises generating the meta-frame based on a reliability of the first event information and a reliability of the second event information being equal to or greater than a first threshold value” claimed in dependent claims 5 and 15, in combination with the remainder of the limitations of the claims, are neither anticipated nor obvious in view of the prior art of record. In the closest prior art found, Fan Chiang (US 2016/0125247 A1), teaches receiving event information from two different event analysis source as shown in Figure 2. PNG media_image1.png 563 751 media_image1.png Greyscale However, Fan Chiang fails to teach generating the meta-frame by comparing the reliability of the event informations from both the first and second image analysis module to a threshold value. Therefore claims 5 and 15 would be allowable for claiming the limitation “receiving first event information from a first event analysis source and receiving second event information from a second event analysis source, and wherein the generating the meta-frame comprises generating the meta-frame based on a reliability of the first event information and a reliability of the second event information being equal to or greater than a first threshold value”. Claims 6 and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The claimed features such as “wherein the generating the meta-frame comprises, in a case where one of a reliability of the first event information and a reliability of the second event information is less than a first threshold value, generating the meta-frame based on the other one of the reliability of the first event information and the reliability of the second event information being equal to or greater than a second threshold value higher than the first threshold value” claimed in dependent claims 6 and 16, in combination with the remainder of the limitations of the claims, are neither anticipated nor obvious in view of the prior art of record. In the closest prior art found, Fan Chiang (US 2016/0125247 A1), teaches receiving event information from two different event analysis source as shown in Figure 2. PNG media_image1.png 563 751 media_image1.png Greyscale However, Fan Chiang fails to teach generating the meta-frame by comparing the reliability of the event informations from both the first and second image analysis module to a first and second threshold value. Therefore claims 6 and 16 would be allowable for claiming the limitation “wherein the generating the meta-frame comprises, in a case where one of a reliability of the first event information and a reliability of the second event information is less than a first threshold value, generating the meta-frame based on the other one of the reliability of the first event information and the reliability of the second event information being equal to or greater than a second threshold value higher than the first threshold value”. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DENISE G ALFONSO whose telephone number is (571)272-1360. The examiner can normally be reached Monday - Friday 7:30 - 5:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DENISE G ALFONSO/Examiner, Art Unit 2662 /AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Apr 19, 2024
Application Filed
Apr 22, 2026
Non-Final Rejection mailed — §103
Jul 20, 2026
Applicant Interview (Telephonic)
Jul 20, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688698
OBSTACLE RECONGNITION METHOD APPLIED TO AUTOMATIC TRAVELING DEVICE AND AUTOMATIC TRAVELING DEVICE
3y 2m to grant Granted Jul 21, 2026
Patent 12682014
CLUSTERING METHOD AND SYSTEM FOR ROAD OBJECT ELEMENTS OF CROWDSOURCED MAP, AND STORAGE MEDIUM
2y 5m to grant Granted Jul 14, 2026
Patent 12657938
A Method for Detecting Correspondence of a Built Structure with a Designed Structure
3y 9m to grant Granted Jun 16, 2026
Patent 12608942
OBSTACLE RECOGNITION METHOD AND APPARATUS, DEVICE, MEDIUM AND WEEDING ROBOT
2y 11m to grant Granted Apr 21, 2026
Patent 12586352
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD AND STORAGE MEDIUM
3y 1m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
89%
With Interview (+14.5%)
2y 12m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 116 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month