DETAILED ACTION
1. This Office Action is responsive to claims filed for App. 18/492,234 on June 15, 2026. Claims 1-15 are pending.
America Invents Act
2. The present application is being examined under the pre-AIA first to invent provisions.
Claim Rejections - 35 USC § 103
3. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
4. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
5. Claim 11 are rejected under 35 U.S.C. 103 as being unpatentable over Oh et al.
( US 2023/0196817 A1 ) in view of Deng et al. ( US 2022/0109885 A1 ).
Oh teaches in Claim 11:
An image segmentation method in a camera device ( [0054] disclose a computing device, such as a smart phone, with an activated camera to capture the digital video 202 ), the method comprising:
receiving, by a neural network ( Figures 2 and 9, [0106] disclose a joint-based segmentation system 106 which has a segmentation neural network 914 ), an image frame including red-green-blue channels ( Figure 2, [0054] discloses a digital video 202, made up of a series of video frames 206, and focuses on an object 204 (read as a region indicative of one or more instances) in the current frame, as detailed in [0057] as well. Please note Figure 3, [0061] which details a current and preceding frame of the digital video as well. [0067] discloses the RGB aspects );
generating, by a template generator, a template including one or more color coded objects from the image frame ( Figures 2 and 3, [0084], [0099] disclose details of the joint-based segmentation system 106 which can predict masks, key point (joints). Please note these prediction aspects/outputs as a prediction template. [0105] discloses the join-based segmentation system 106 modifies the digital video by adding text, color, or other visual effects to the digital video. As for instances, please not details above of the object 204 ); and
merging, by a template encoder, a template including the one or more color coded objects with the red-green-blue channels of image frames subsequent to the image frame as a preprocessed input for image segmentation in the neural network ( Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes )
Oh does not explicitly teach “adding, by a template encoder, respective predetermined fractions of a template including the one or more color coded objects and the red-green-blue channels of image frames subsequent to the image frame as a preprocessed input for image segmentation in the neural network”.
Initially, Oh teaches: Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes. Furthermore, Oh teaches in [0098] of temporal masks and in general, encoding.
To emphasize, in the same field of endeavor, frame encoding, Deng teaches of encoding multiple consecutive (subsequent) frames, ( Deng, Figure 4, [0280] ). Notably, luma/chroma, which is for color aspects, can be incorporated in block partitions (read as fractional). The goal for Deng is to construct a motion vector and [0229] discloses a fractional motion vector may be used to derive the corresponding luma (color) samples to result in an interpolation process. Furthermore, Deng teaches in [0635] of using weighted averages of color components in the reference frames (another indication of a fractional aspect). As combined, in order to derive subsequent frames, fractional amounts can be used.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the fractional aspects, as taught by Deng, with the motivation that an interpolation process can be used to better derive the subsequent frames, ( Deng, [0229] ).
6. Claims 1-10 and 12-15 are rejected under 35 U.S.C. 103 as being unpatentable over Oh et al. ( US 2023/0196817 A1 ) in view of Deng et al. ( US 2022/0109885 A1 ) and Wang et al. ( US 2020/0294240 A1 ).
Oh teaches in Claim 1:
A method for encoding temporal information in an electronic device, the method comprising:
identifying, by a neural network, at least one region indicative of one or more objects in a first frame by analyzing a first frame among a plurality of frames ( Figure 2, [0054] discloses a digital video 202, made up of a series of video frames 206, and focuses on an object 204 (read as a region indicative of one or more objects) in the current frame, as detailed in [0057] as well. Please note Figure 3, [0061] which details a current and preceding frame of the digital video as well );
outputting, by the neural network ( Figures 2 and 9, [0106] disclose a joint-based segmentation system 106 which has a segmentation neural network 914 ), a first prediction template for the first frame including the one or more objects in the first frame ( Figures 2 and 3, [0084], [0099] disclose details of the joint-based segmentation system 106 which can predict masks, key point (joints). Please note these prediction aspects/outputs as a prediction template );
generating, by a template generator, a color coded template for the first frame by assigning a color to each of the one or more objects included in the first prediction template ( [0105] discloses the join-based segmentation system 106 modifies the digital video by adding text, color, or other visual effects to the digital video ); but
Oh does not explicitly teach “generating, by a template encoder, a temporal encoded second frame by adding respective predetermined fractions of the first color coded template and a second frame subsequent to the first frame among the plurality of frames”.
Initially, Oh teaches: Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes. Furthermore, Oh teaches in [0098] of temporal masks and in general, encoding.
To emphasize, in the same field of endeavor, frame encoding, Deng teaches of encoding multiple consecutive (subsequent) frames, ( Deng, Figure 4, [0280] ). Notably, luma/chroma, which is for color aspects, can be incorporated in block partitions (read as fractional). The goal for Deng is to construct a motion vector and [0229] discloses a fractional motion vector may be used to derive the corresponding luma (color) samples to result in an interpolation process. Furthermore, Deng teaches in [0635] of using weighted averages of color components in the reference frames (another indication of a fractional aspect). As combined, in order to derive subsequent frames, fractional amounts can be used.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the fractional aspects, as taught by Deng, with the motivation that an interpolation process can be used to better derive the subsequent frames, ( Deng, [0229] ).
Oh does not explicitly teach of “supplying a temporal encoded second frame to the neural network for generating a second prediction template for the second frame”.
However, in the same field of endeavor, segmentation models, Wang teaches of a first and second prediction module 401 and 402, respectively, ( Wang, Figure 5, [0068]-[0072] ). Notably, Wang teaches in [0072] that the second prediction module 402 inputs the data from the first prediction module and mask parameters of the second-category objects can be predicted. To clarify, Oh teaches of an initial prediction aspect and Wang teaches of a second prediction template for the same frame modified by the first prediction template, i.e. the interpreted second frame of Oh. It is the same frame that is further analyzed in this situation.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the second prediction module, as taught by Wang, with the motivation that by using both prediction modules, notably the second prediction module, second-category objects can be segmented, further enhancing the process, ( Wang, [0074] ).
Oh and Deng teach in Claim 2:
The method of claim 1, further comprising: identifying, by the neural network, at least one region indicative of one or more objects in the temporal encoded second frame by analyzing the temporal encoded second frame; outputting, by the neural network, the second prediction template including the one or more objects in the temporal encoded second frame; generating, by the template generator, a second color coded template for the temporal encoded second frame by applying at least one color to the second prediction template; generating, by the template encoder, a temporal encoded third frame, by combining a third frame and the second color coded template; and supplying the temporal encoded third frame to the neural network. ( This claim introduces a third frame and essentially repeats the steps found in Claim 1, but between a second and third frame, as opposed to a first and second frame. As combined, the combined teachings of Oh can be performed for multiple (at least three) frames )
Oh teaches in Claim 3:
The method of claim 1, wherein the plurality of frames is from a preview of a capturing device, and wherein the plurality of frames is represented by a red-green-blue (RGB) color model. ( Figure 3, [0067] discloses RGB frame aspects )
As per Claim 4:
Oh may not explicitly teach “wherein the fraction of the first color coded template is 0.1.”
Oh teaches in [0027] of modifying the digital video using the joint-based segmentation masks and [0105] discloses such modifications include color aspects, etc. Furthermore, please note the combination with Deng to use a fractional motion vector analysis.
Respectfully, it is clear that the modification of the frame data needs to be handled smoothly and avoiding segmentation inaccuracies, [0003]. The specific blending fraction value is a design choice/optimization issue as one of ordinary skill in the art would realize the optimized value would result in a more accurate modification when modifying the digital video.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the proper blending fraction, with the motivation that this is a design choice/optimization issue to result in better accuracy and modification of frame data.
Oh teaches in Claim 5:
The method of claim 1, wherein the neural network is one of a segmentation neural network or an object detection neural network. ( [0056] discloses the joint-based segmentation system 106 includes a segmentation neural network )
Oh teaches in Claim 6:
The method of claim 5, wherein output of the segmentation neural network includes one or more segmentation masks of the one or more objects in the first frame. ( [0057] discloses masks for each frame o the digital video that portrays the object )
Oh teaches in Claim 7:
The method of claim 5, wherein output of the object detection neural network includes one or more bounding boxes of the one or more objects in the first frame. ( [0064] discloses the joint-based segmentation system 106 generate a bounding box, among other aspects for objects portrayed in the frame )
Oh teaches in Claim 8:
The method of claim 1, wherein the electronic device includes a smartphone or a wearable device that is equipped with a camera. ( [0054] disclose a computing device, such as a smart phone, with an activated camera to capture the digital video 202 )
Oh teaches in Claim 9:
The method of claim 1, wherein the neural network is configured to receive the first frame prior to analyzing the first frame. ( Figure 2 discloses the joint-based segmentation system 106 receives the digital video and its plurality of video frames and then performs analysis )
Oh teaches in Claim 10:
An intelligent instance segmentation method in a device, the method comprising:
receiving, by a neural network, a first frame from among a plurality of frames ( Figure 2, [0054] discloses a digital video 202, made up of a series of video frames 206, and focuses on an object 204 (read as a region indicative of one or more objects, as detailed below) in the current frame, as detailed in [0057] as well. Please note Figure 3, [0061] which details a current and preceding frame of the digital video as well );
analyzing, by the neural network ( Figures 2 and 9, [0106] disclose a joint-based segmentation system 106 which has a segmentation neural network 914 ), the first frame to identify a region indicative of one or more objects in the first frame; generating, by the neural network, a template having the one or more instances in the first frame ( Figures 2 and 3, [0084], [0099] disclose details of the joint-based segmentation system 106 which can predict masks, key point (joints). Please note these prediction aspects/outputs as a prediction template );
assigning, by a template generator, a color to each of the one or more objects included in the first template ( [0105] discloses the join-based segmentation system 106 modifies the digital video by adding text, color, or other visual effects to the digital video );
receiving, by the neural network, a second frame; generating, by a template encoder, a temporal encoded second frame by merging the color coded template with the second frame ( Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes. This modification is merging the frame with the color aspects, etc. Furthermore, Oh teaches in [0098] of temporal masks and in general, encoding ); but
Oh does not explicitly teach of receiving a second frame “subsequent to the first frame” and generating, by a template encoder, a temporal encoded second frame by adding respective predetermined fractions of the color coded template and the second frame”.
Initially, Oh teaches: Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes. Furthermore, Oh teaches in [0098] of temporal masks and in general, encoding.
Initially, Oh teaches: Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes. Furthermore, Oh teaches in [0098] of temporal masks and in general, encoding.
To emphasize, in the same field of endeavor, frame encoding, Deng teaches of encoding multiple consecutive (subsequent) frames, ( Deng, Figure 4, [0280] ). Notably, luma/chroma, which is for color aspects, can be incorporated in block partitions (read as fractional). The goal for Deng is to construct a motion vector and [0229] discloses a fractional motion vector may be used to derive the corresponding luma (color) samples to result in an interpolation process. Furthermore, Deng teaches in [0635] of using weighted averages of color components in the reference frames (another indication of a fractional aspect). As combined, in order to derive subsequent frames, fractional amounts can be used.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the fractional aspects, as taught by Deng, with the motivation that an interpolation process can be used to better derive the subsequent frames, ( Deng, [0229] ).
Oh does not explicitly teach “for generating a second template for segmenting the one or more objects in the temporal encoded second frame”.
However, in the same field of endeavor, segmentation models, Wang teaches of a first and second prediction module 401 and 402, respectively, ( Wang, Figure 5, [0068]-[0072] ). Notably, Wang teaches in [0072] that the second prediction module 402 inputs the data from the first prediction module and mask parameters of the second-category objects can be predicted. To clarify, Oh teaches of an initial prediction aspect and Wang teaches of a second prediction template for the same frame modified by the first prediction template, i.e. the interpreted second frame of Oh. It is the same frame that is further analyzed in this situation.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the second prediction module, as taught by Wang, with the motivation that by using both prediction modules, notably the second prediction module, second-category objects can be segmented, further enhancing the process, ( Wang, [0074] ).
Oh teaches in Claim 12:
A system for encoding temporal information, comprising:
a capturing device including a camera ( [0054] disclose a computing device, such as a smart phone, with an activated camera to capture the digital video 202 );
a neural network ( Figures 2 and 9, [0106] disclose a joint-based segmentation system 106 which has a segmentation neural network 914 ), wherein the neural network is configured to:
identify at least one region indicative of one or more objects in a first frame by analyzing the first frame among a plurality of frames from the capturing device ( Figure 2, [0054] discloses a digital video 202, made up of a series of video frames 206, and focuses on an object 204 (read as a region indicative of one or more instances) in the current frame, as detailed in [0057] as well. Please note Figure 3, [0061] which details a current and preceding frame of the digital video as well. As noted above, the camera provides the digital video ), and
output a first prediction template for the first frame including the one or more objects in the first frame ( Figures 2 and 3, [0084], [0099] disclose details of the joint-based segmentation system 106 which can predict masks, key point (joints). Please note these prediction aspects/outputs as a prediction template ), and
a template generator configured to generate a first color coded template for the first frame by assigning a color to each of the one or more objects included in the first prediction template ( [0105] discloses the join-based segmentation system 106 modifies the digital video by adding text, color, or other visual effects to the digital video ); and
a template encoder configured to generate a temporal encoded second frame by merging a second frame and the first color coded template ( Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes. Furthermore, Oh teaches in [0098] of temporal masks and in general, encoding. Also, please note the combination below as well ); but
Oh does not explicitly teach “generate a temporal encoded second frame by adding respective predetermined fractions of the first color coded template an a second frame subsequent to the first frame”.
Initially, Oh teaches: Figure 2, [0059] discloses the joint-based segmentation system 106 modifies the digital video 202 with the additions noted above, those additions including the possibility of color changes. Furthermore, Oh teaches in [0098] of temporal masks and in general, encoding.
To emphasize, in the same field of endeavor, frame encoding, Deng teaches of encoding multiple consecutive (subsequent) frames, ( Deng, Figure 4, [0280] ). Notably, luma/chroma, which is for color aspects, can be incorporated in block partitions (read as fractional). The goal for Deng is to construct a motion vector and [0229] discloses a fractional motion vector may be used to derive the corresponding luma (color) samples to result in an interpolation process. Furthermore, Deng teaches in [0635] of using weighted averages of color components in the reference frames (another indication of a fractional aspect). As combined, in order to derive subsequent frames, fractional amounts can be used.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the fractional aspects, as taught by Deng, with the motivation that an interpolation process can be used to better derive the subsequent frames, ( Deng, [0229] ).
Oh may not explicitly teach to “supply the temporal encoded second frame to the neural network for generating a second prediction template for the second frame”.
However, in the same field of endeavor, segmentation models, Wang teaches of a first and second prediction module 401 and 402, respectively, ( Wang, Figure 5, [0068]-[0072] ). Notably, Wang teaches in [0072] that the second prediction module 402 inputs the data from the first prediction module and mask parameters of the second-category objects can be predicted. To clarify, Oh teaches of an initial prediction aspect and Wang teaches of a second prediction template for the same frame modified by the first prediction template, i.e. the interpreted second frame of Oh. It is the same frame that is further analyzed in this situation.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the second prediction module, as taught by Wang, with the motivation that by using both prediction modules, notably the second prediction module, second-category objects can be segmented, further enhancing the process, ( Wang, [0074] ).
Oh teaches in Claim 13:
The system of claim 12, wherein the neural network is configured to receive the first frame. ( [0059] discloses the joint-based segmentation system 106 receives the frame data in order to analyze and modify )
Oh teaches in Claim 14:
The system of claim 12, wherein the plurality of frames of the capturing device is represented by a red-green-blue (RGB) color model. ( Figure 3, [0067] discloses RGB frame aspects )
As per Claim 15:
Oh may not explicitly teach “wherein the fraction of the first color coded template value is 0.1.”
Oh teaches in [0027] of modifying the digital video using the joint-based segmentation masks and [0105] discloses such modifications include color aspects, etc. Also, please note the combination with Deng with the fractional motion vector aspects.
Respectfully, it is clear that the modification of the frame data needs to be handled smoothly and avoiding segmentation inaccuracies, [0003]. The specific blending fraction value is a design choice/optimization issue as one of ordinary skill in the art would realize the optimized value would result in a more accurate modification when modifying the digital video.
Therefore, it would have been obvious to one of ordinary skill in the art, at the effective filed date of the invention, to implement the proper blending fraction, with the motivation that this is a design choice/optimization issue to result in better accuracy and modification of frame data.
Response to Arguments
7. Applicant’s arguments considered, but are respectfully moot in view of new grounds of rejection(s).
Please note the updated rejection in light of the claim amendments, notably the reliance on the Deng reference. As a result, Applicant’s arguments are moot at this time.
Conclusion
8. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DENNIS P JOSEPH whose telephone number is (571)270-1459. The examiner can normally be reached Monday - Friday 5:30 - 3:30 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amr Awad can be reached at 571-272-7764. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DENNIS P JOSEPH/Primary Examiner, Art Unit 2621