DETAILED ACTION
1. This communication is being filed in response to the submission having a mailing date of 06/05/2026 in which a (3) month Shortened Statutory Period for Response has been set.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Acknowledgements
3. Upon new entry, claims (1 -20) remain pending on this application, of which (1, 8, 15) being the three (3) parallel running independent claims on record.
Examiner thanks’ Applicant representative (Atty. C. Luk; Reg. No, 73,991) for the detailed remarks and clarifications, and for the cooperation expediting the case.
The previous rejection under 35 USC 103 is maintained, as no persuasive arguments presented. See “Response to Arguments” section below, for more details.
Information Disclosure Statement
4. It is note that no information disclosure statements (IDS) were filed with this application. However, the Background & introduction Sections of Applicant's Specification describe several known materials and technologies that seem pertinent to be presented by the Applicant, moving forward.
Response to arguments
5. Applicant’s arguments have been carefully considered, but they’re not persuasive, for at least the following reasons:
5.1. For record clarity, the following terms/limitations will be read as following:
_ "auxiliar facial signals" as - (e.g. facial motion information (e.g., expression and head pose) and complex background changes in a scalable and flexible manner; [specs; 0083]).
_ “auxiliar generative synthesis model” as (e.g. synthesis model 512 configures one or more processors of a computing system to reconstruct the face video sequence based on the reconstructed compact representation [0090]).
5.2. The undersigned considers that the no allowable subject matter has been yet identified in the claims. The claims language of the three (3) parallel running independent claims (1, 8, 15) and the associated dependencies, are drawn to - a codec ecosystem (encoder/decoder) of the same, in accordance with the HEVC/VVC/AV1 and legacy codec standards, employing the commonly known GFVC technique, for high quality image/face reconstruction, similarly and detailed described in the applied combination of PA (AAPA/Rider). Applicant further acknowledged the use of deep learning AI in codec algorithms techniques, improving quality and codec efficiency in at least [0002].
Rider similarly teaches: (e.g. an Intelligence based computer system that employes generative modeling [10: 06], for a codec implementations (i.e. encoder 110, and decoder 120, and CRM storage; [Rider; 1: 15]) for video/image(s) processing, comprising: identify, process and analyze highly sensitive visual regions (i.e. face, eye, mouth, etc; Figs. (4 and 7); [Rider; 5: 01]), inputted to first and second channels/models, as shown in Figs (2, 5, 9, 10) [Rider; Cols. 9 -10]; using codec quantization feature seps; [Rider; Col. 9]; measuring and estimating the canonical error difference [Rider; Col. 10]; boosting the quality of the reconstructed image in the process; [Col. 10: 65], in accordance with the codec and legacy standards [Rider; 5: 33].
5.3. Applicant argues the following topics:
_ a failure to disclose [“facial signals”]; Examiner disagrees because under the broadest reasonable BRI interpretation, consistent with the instant Specs, and the abilities of one skilled in the art, at least the AAPA discloses the use of "facial characterization data parameters (including "face motion") at the pixel level" [AAPA; 0002]. Rider analogously teaches the same in at least [Cols. 11-12].
_ a failure to disclose [“auxiliar generative synthesis model “]; Examiner disagrees because under the same BRI interpretation, at least the AAPA discloses the use of "synthesis model for image reconstruction" in at least [AAPA; 0002]. Rider analogously teaches - inputted image data into models, as shown in Figs (2, 5, 9, 10) [Rider; Cols. 9 -10];
_ a failure to disclose [reconstructing a plurality of reconstructed inter frames… page 9]; – Examiner respectfully disagrees, because under the broadest reasonable BRI interpretation, consistent with the instant Specs, and the abilities of one skilled in the art, the reconstruction step of codec process, by definition implies inputting/processing intra/spatial and inter/motion frame prediction data, emphasis added.
_ a failure to disclose [extracting an original auxiliary facial signal from plurality of original inter frames,"… page 9]; – Examiner respectfully disagrees, because under the same BRI, at least AAPA discloses facial motion (auxiliar facial signal), similarly inputted in the model; [AAPA]. Rider similarly discloses extracting “auxiliar pixel information” inputted to AI training model, in at least Figs. (3-5); [Cols. 11-12].
_ a failure to disclose [boosting generation quality of the reconstructed inter frames "… page 10]; – Examiner respectfully disagrees, because under the same BRI, at least AAPA discloses "high quality image reconstruction using the model" [0002], prioritizing the quality of the generated content, and efficient representation and reconstruction of the original video; [0003]. Rider similarly discloses boosting the quality of the reconstructed image in the process; [Col. 10: 65], in accordance with the codec and legacy standards [Rider; 5: 33].
5.4. It is valid to point out that in order to prove patentability at USPTO, the claim language must present a clear defined functionality, and an algorithm execution, that differentiates from the common knowledge, that would produce a certain effect and/or result, by executing a series of acts/steps, able to transform and reduce them to a different state of thing. The presented list of claims, (as currently stated) fails this requirement.
5.5. Examiner also notes that Applicant lists plurality of well-known techniques following the passive term(s) such as “reconstructing/extracting/boosting …etc” that passively indicates that a function is performed without requiring the/any functional structure and/or method steps, as a limitation on the claim itself. It is clear that such claim language does not further limit the claims, and does not require a separate reason for rejection; (see MPEP 2111.04). The clause may be given some weight to the extent it provides "meaning and purpose” to the claimed invention, but not when “it simply expresses the intended result” of the invention.
5.6. Regarding the rationale and motivation for the associated features/steps presented in the claims, please refer to Rejection section (6) for details.
Finally, the Office considers Applicant's arguments not persuasive, as applied rejection on record as a whole reads on the claimed construction, establishing the "Prima Facie" case of equivalent disclosures, on the basis of a person of ordinary skills in the art would have recognized the similar elements shown, or the same structural similarities shown, wherein such structure/methodology performs the same identical functions in substantially the same way, able to produce the same identical results.
_ See MPEP – 2183; making the Prima Facie Case of Equivalence.
_ See In re Bond, 910 F.2d 831, 833, 15 USPQ2d 1566]; when similar structure applies;
_ See Kemco Sales, Inc. vs. Control Papers., 208 F.3d 1352, 54 USPQ2d 1308]; when identical functionality is specified in the claim, in substantially the same way.
Claim rejection section
35 USC 103
6. In the event the determination of the status of the application as subject to AIA 35
U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6.1. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ
459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
6.2. Claims (1 -20) is/are rejected under 35 U.S.C. 103 as being unpatentable over the Applicant’s admitted Prior Art (AAPA), hereafter “AAPA”), in view of Rider; et al. (US 11,532,104 B2; hereafter “Rider”).
Claim 1. AAPA discloses the principles of the invention substantially as claimed - A computing system, comprising: (e.g. a codec ecosystem (i.e. encoder and decoder) of the same [AAPA]); in complain with “earlier” common HEVC/VVC codec format(s); [0015]
one or more processors, and (e.g. see analogous standard encoder/decoder components disclosed);
a computer-readable storage medium communicatively coupled to the one or more processors, the computer-readable storage medium storing computer-readable instructions executable by the one or more processors that, when executed by the one or more processors, perform associated operations comprising: (e.g. see analogous standard encoder/decoder components and algorithm techniques, disclosed; [AAPA]);
reconstructing a plurality of reconstructed inter frames of a video sequence by inputting a plurality of original inter frames (e.g. decoder for reconstruction interframe (i.e. temporal (i.e. motion, inter) facial data; [AAPA; 0002])
to a generative face video compression (“GFVC”) model; (e.g. using GFVC compression technique, employing synthesis modeling of pixel level facial representations (i.e. 2D/3D key-points, temporal/inter feature(s), segmentation map(s), facial semantics, etc); [AAPA; 0002]);
extracting an original auxiliary facial signal from the plurality of original inter frames; (e.g. see original pixel-level facial signal representation of the same; [AAPA; 0002]);
extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames; (e.g. see auxiliar generative synthesis model disclosed; [AAPA; 0003])
predicting a reconstructed auxiliary facial signal from a quantized auxiliary facial signal, (e.g. see image/segment reconstruction; [AAPA; 0003]);
based on a difference between the model-generated auxiliary facial signal and the original auxiliary facial signal; (e.g. similar analysis and synthesis step(s); [AAPA; 0002])
and boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal; (e.g. enhancing visual rich features using modeling; [AAPA; 0003]
It is note that AAPA discloses the status of the Art, at the time of the invention in general terms, and therefore some specifics components are missed and//or not fully described.
For the purpose of additional clarification and in the same field of endeavor, Rider teaches: (e.g. an Intelligence based computer system that employes generative modeling [10: 06], for lossy codec implementations (i.e. encoder 110, and decoder 120, and CRM storage; [Rider; 1: 15]) for video/image(s) processing, comprising: identify, process and analyze highly sensitive visual regions (i.e. face, eye, mouth, etc; Figs. (4 and 7); [Rider; 5: 01]), inputted to first and second channels/models, as shown in Figs (2, 5, 9, 10) [Rider; Cols. 9 -10]; using codec quantization feature seps; [Rider; Col. 9]; measuring and estimating the canonical error difference [Rider; Col. 10]; and boosting the quality of the reconstructed image in the process; [Col. 10: 65].
Therefore, it would have been obvious to one skilled in the art before the effective filing date of the claimed invention, to modify the teachings of (AAPA) section, with the codec system implementation of Rider, in order to similarly provide – (e.g. coding efficiency by optimize the difference bw. original data-image and modeled data [1: 45], reducing the canonical error distortion, and simultaneously reducing the bandwidth requirements [Rider; 1: 50; 9: 42].
Claim 2. AAPA/Rider discloses - The computing system of claim 1, wherein extracting the original auxiliary facial signal from the plurality of original inter frames comprises: (e.g. see inputted to first and second channels/models, Figs (5, 9, 10) [Rider; Col. 9 -10; 12: 01];
downsampling the plurality of original inter frames; and transforming the plurality of original inter frames to a high-dimensional face feature map; (e.g. image data resampling (X, Xf), using multipliers, in Figs. (5, 9, 10); [Rider; 14: 20]);
and extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises: downsampling the plurality of reconstructed inter frames; (5, 9, 10); [Rider; 14: 20]); and transforming the plurality of reconstructed inter frames to a high-dimensional face feature map; (e.g. see combined (fo, f’o) for higher bit allocation, Figs. (5, 9, 10); [14: 30]; the same motivation applies herein.)
Claim 3. AAPA/Rider discloses - The computing system of claim 1, wherein extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises: (e.g. resampling (X, Xf), using multipliers, in Figs. (5, 9, 10); [Rider; 14: 20]);
selecting a higher or lower granularity of the original auxiliary facial signal based on higher or lower bitstream bandwidth; (e.g. able to reduce the signal granularity (i.e. canonical error distortion), and simultaneously minimizing bandwidth requirements [1: 50; 9: 42]; the same motivation applies herein.)
Claim 4. AAPA/Rider discloses - The computing system of claim 1, wherein the reconstructed auxiliary facial signal is predicted based further on a Gaussian distribution comprising entropy parameters. Examiner’s note is taken: In probability theory and statistics, a Gaussian process is a stochastic process (a collection of random variables indexed by time or space), such that every finite collection of those random variables has a multivariate normal distribution. In addition, Rider - (e.g. employes training modeling using similar gradient approach; [Rider; 13: 25]; same motivation applies herein.
Claim 5. AAPA/Rider discloses - The computing system of claim 4, wherein the entropy parameters are conditioned upon: a hyperprior comprising the model-generated auxiliary facial signal; and a causal context of the quantized auxiliary facial signal; (e.g. model generation using quantized data form the codec implementation; [Rider; Col. 9]; the same motivation applies herein.)
Claim 6. AAPA/Rider discloses - The computing system of claim 1, wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal comprises: (e.g. see boosting the quality of the reconstructed image in the process; [Col. 10: 65].
transforming the reconstructed inter frames into facial features; (e.g. see data reconstruction of the same in Figs. (5, 9, 10); [Rider]);
transforming the reconstructed auxiliary facial signal into signal features having a same feature dimensionality as the facial features; (e.g. see data reconstruction of the same in Figs. (5, 9, 10) producing the original features; [Rider; Col. 10])
performing linear projection upon the facial features and the signal features to yield latent feature maps; (e.g. see produce latent representation; [Rider; Cols. 9 -10])
and inputting the latent feature maps into an attention layer to yield fused attention features; (e.g. see merging/fusion in Figs. (5, 9, 10); [Rider; Col. 14]; the same motivation applies herein.)
Claim 7. AAPA/Rider discloses - The computing system of claim 6, wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal further comprises: (e.g. see image reconstructed image in Figs. (5, 9, 10); [Rider; Cols. 9-10].
inputting the attention features to a coarse face generator U-Net decoder to yield coarsely enhanced inter frames; (e.g. U-net analogous codecs (i.e. fully convolutional codec operational, disclosed; [Rider; 6: 26]);
learning a motion estimation field and a facial occlusion map by concatenating a reconstructed key-reference frame and the coarsely enhanced inter frames; (e.g. see merging/fusion in Figs. (5, 9, 10); [Rider; Col. 14]);
and applying the motion estimation field and the facial occlusion map to multi-scale spatial features derived from the reconstructed key-reference frame; (e.g. see prediction parameters (i.e. intra/inter prediction) and multi scaling/sampling, similarly used; [Rider; 9: 50]; the same motivation applies herein.)
Claim 8. AAPA/Rider discloses - A method, comprising:
reconstructing a plurality of reconstructed inter frames of a video sequence by inputting a plurality of original inter frames to a generative face video compression (“GFVC”) model; extracting an original auxiliary facial signal from the plurality of original inter frames;
extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames; predicting a reconstructed auxiliary facial signal based on a difference between the model-generated auxiliary facial signal and the original auxiliary facial signal; and boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal. (Current lists all the same elements as recite in Claim 1 above, but in “Method form” instead, and is/are therefore on the same premise.)
Claim 9. AAPA/Rider discloses - The method of claim 8, wherein extracting the original auxiliary facial signal from the plurality of original inter frames comprises: downsampling the plurality of original inter frames; and transforming the plurality of original inter frames to a high-dimensional face feature map; and extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises: downsampling the plurality of reconstructed inter frames; and transforming the plurality of reconstructed inter frames to a high-dimensional face feature map. (The same rationale and motivation apply as given to Claim (2) above.)
Claim 10. AAPA/Rider discloses - The method of claim 8, wherein extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises: selecting a higher or lower granularity of the original auxiliary facial signal based on higher or lower bitstream bandwidth. (The same rationale and motivation apply as given to Claim (3) above.)
Claim 11. AAPA/Rider discloses - The method of claim 8, wherein the reconstructed auxiliary facial signal is predicted based further on a Gaussian distribution comprising entropy parameters. (The same rationale and motivation apply as given to Claim (4) above.)
Claim 12. AAPA/Rider discloses - The method of claim 11, wherein the entropy parameters are conditioned upon: a hyperprior comprising the model-generated auxiliary facial signal; and a causal context of the quantized auxiliary facial signal. (The same rationale and motivation apply as given to Claim (5) above.)
Claim 13. AAPA/Rider discloses - The method of claim 8, wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal comprises: transforming the reconstructed inter frames into facial features; transforming the reconstructed auxiliary facial signal into signal features having a same feature dimensionality as the facial features; performing linear projection upon the facial features and the signal features to yield latent feature maps; and inputting the latent feature maps into an attention layer to yield fused attention features. (The same rationale and motivation apply as given to Claim (6) above.)
Claim 14. AAPA/Rider discloses - The method of claim 13, wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal further comprises: inputting the attention features to a coarse face generator U-Net decoder to yield coarsely enhanced inter frames; learning a motion estimation field and a facial occlusion map by concatenating a reconstructed key-reference frame and the coarsely enhanced inter frames; and applying the motion estimation field and the facial occlusion map to multi-scale spatial features derived from the reconstructed key-reference frame. (The same rationale and motivation apply as given to Claim (7) above.)
Claim 15. AAPA/Rider discloses - One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform operations comprising: reconstructing a plurality of reconstructed inter frames of a video sequence by inputting a plurality of original inter frames to a generative face video compression (“GFVC”) model; extracting an original auxiliary facial signal from the plurality of original inter frames; extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames; predicting a reconstructed auxiliary facial signal based on a difference between the model-generated auxiliary facial signal and the original auxiliary facial signal; and boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal. (Current lists all the same elements as recite in Claim 1 above, but in “CRM form” instead, and is/are therefore on the same premise.)
Claim 16. AAPA/Rider discloses - The non-transitory computer-readable media of claim 15, wherein extracting the original auxiliary facial signal from the plurality of original inter frames comprises:
downsampling the plurality of original inter frames; and transforming the plurality of original inter frames to a high-dimensional face feature map; and extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises:
downsampling the plurality of reconstructed inter frames; and transforming the plurality of reconstructed inter frames to a high-dimensional face feature map. (The same rationale and motivation apply as given to Claim (2) above.)
Claim 17. AAPA/Rider discloses - The non-transitory computer-readable media of claim 15, wherein extracting a model-generated auxiliary facial signal from the plurality of reconstructed inter frames comprises: selecting a higher or lower granularity of the original auxiliary facial signal based on higher or lower bitstream bandwidth. (The same rationale and motivation apply as given to Claim (3) above.)
Claim 18. AAPA/Rider discloses - The non-transitory computer-readable media of claim 15, wherein the reconstructed auxiliary facial signal is predicted based further on a Gaussian distribution comprising entropy parameters. (The same rationale and motivation apply as given to Claim (4) above.)
Claim 19. AAPA/Rider discloses - The non-transitory computer-readable media of claim 15, wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal comprises: transforming the reconstructed inter frames into facial features; transforming the reconstructed auxiliary facial signal into signal features having a same feature dimensionality as the facial features; performing linear projection upon the facial features and the signal features to yield latent feature maps; and inputting the latent feature maps into an attention layer to yield fused attention features. (The same rationale and motivation apply as given to Claim (6) above.)
Claim 20. AAPA/Rider discloses - The non-transitory computer-readable media of claim 19, wherein boosting generation quality of the reconstructed inter frames based on the reconstructed auxiliary facial signal further comprises: inputting the attention features to a coarse face generator U-Net decoder to yield coarsely enhanced inter frames; learning a motion estimation field and a facial occlusion map by concatenating a reconstructed key-reference frame and the coarsely enhanced inter frames; and applying the motion estimation field and the facial occlusion map to multi-scale spatial features derived from the reconstructed key-reference frame. (The same rationale and motivation apply as given to Claim (7) above.)
Prior Art Citations
7. The following List of prior art, made of record and not relied upon, is/are considered pertinent to applicant's disclosure:
7.1. Patent documentation:
US 12,621,492 B2 H04N19/70; H04N19/124; H04N19/136; Chen; Bolin et al.
US 12,470,746 B2 H04N19/172; H04N19/597; H04N7/147; Chen; Bolin et al.
US 12,477,120 B2 G06T3/18; G06T3/20; G06T3/60; Wang; Zhao et al.
US 11,532,104 B2 G06N3/045; G06N3/0455; G06N3/0464; Ryder; et al.
7.2. Non-Patent documentation:
_ First order motion model for image animation; 2019.
_ The Unreasonable Effectiveness of Deep Features as a Perceptual Metric; Zhang – 2018.
_ Generative Adversarial Networks for Extreme Learned Image Compression; 2019.
_ One-shot free-view neural talking-head synthesis for video conferencing; Wang – 2020.
_ Beyond key-point coding; Chen – 2022.
_ Interactive face video coding; A generative compression framework; Chen - Febr-2023.
_ Revisiting multimodal representation in Contrastive learning; Chen - March-2023
_ Generative Face Video Coding Techniques and Standardization efforts; Chen - Nov-2023.
_ Pleno-Generation A Scalable GFVC Framework with Bandwidth Intelligence; 2025
_ Beyond GFVC - A Progressive Face Video Compression Framework with Adaptive Visual Tokens; Oct-2024.
CONCLUSIONS
8. In view of the above Examiner’s considerations, THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.1 36(a). See also See MPEP 5 706.07(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
9. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LUIS PEREZ-FUENTES (luis.perez-fuentes@uspto.gov) whose telephone number is (571) 270 -1168. The examiner can normally be reached on Monday-Friday 8am-5pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, WILLIAM VAUGHN can be reached on (571) 272-3922. The fax phone number for the organization where this application or proceeding is assigned is (571) 272 -3922. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated system, please call (800) 786 -9199 (USA OR CANADA) or (571) 272 -1000.
/LUIS PEREZ-FUENTES/
Primary Examiner, Art Unit 2481.