Prosecution Insights
Last updated: October 01, 2026
Application No. 19/050,152

SYSTEMS AND TECHNIQUES FOR RETRAINING MODELS FOR VIDEO QUALITY ASSESSMENT AND FOR TRANSCODING USING THE RETRAINED MODELS

Non-Final OA §103§112
Filed
Feb 11, 2025
Priority
Nov 26, 2019 — nonprovisional of PCTUS2019063191 +1 more
Examiner
DARDANO, STEFANO ANTHONY
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
73 granted / 93 resolved
+18.5% vs TC avg
Strong +32% interview lift
Without
With
+31.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
15 currently pending
Career history
104
Total Applications
across all art units

Statute-Specific Performance

§101
9.9%
-30.1% vs TC avg
§103
57.9%
+17.9% vs TC avg
§102
17.9%
-22.1% vs TC avg
§112
12.6%
-27.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 93 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Status Claims 1-20 are pending. Priority This application is a continuation of U.S. Application Serial No. 17/762,289, filed March 21, 2022, which is a national stage application under 35 U.S.C. 371 of PCT/US2019/063191, filed November 26, 2019. Information Disclosure Statement The IDS filed 02/11/25 has been considered. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 10 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. In claim 10, the limitation “teaches wherein the first user generated content is the second user generated content” is written in a way that renders the scope of the claim unclear. In particular, it is unclear what exactly “the first user generated content is the second user generated content” means. Does it mean the first user generated content is the second user generated content by way of the datasets containing the same type of data (both are user generated images), or does it mean that the first and the second are the same by way of them being physically the same (while they are labeled the first and second user generated content, they actually refer to the same singular dataset and not separate ones). Clarification with support is recommended. For examination purposes the first embodiment (where they are both the same type of data) will be considered the interpreted language for search and finding applicable prior art. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 8-14, 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Talebi et al. (“Learned Perceptual Image Enhancement” Hereinafter “Talebi”) in view of Shan et al. (“Two-Stage Transfer Learning of End-to-End Convolutional Neural Networks for Webpage Saliency Prediction” Hereinafter “Shan”) in further view of Niu et al. (“Siamese-Network-Based Learning to Rank for No-Reference 2D and 3D Image Quality Assessment” Hereinafter “Niu”) in further view of Wang et al. (US 20190122115 A1 Hereinafter “Wang”). Rejection overview: This overview will provide context for the combination and mapping. Talebi is relied upon for the idea that a model originally trained for object detection can be retrained using transfer learning for image quality analysis. Shan describes a two-stage transfer learning, “In fact, the two-stage transfer learning can be regarded as a task transfer in the first stage and a domain transfer in the second stage, respectively. The experimental results indicate that the proposed two-stage transfer learning of end-to-end CNN can obtain a substantial performance improvement for webpage saliency prediction”. Using this two-step method implies improved performance for task changing using transfer learning. Implementing these teachings in Talebi would result in a second model being produced by retaining the first model before it is finetuned to create the third model. The second model would have a task similar to the third model since the second finetuning simply implements a domain change as defined by Shan, and since Talebi’s third model is trained for image quality analysis, the second model would have to be trained for the same task, which would include the claimed “technical content assessment” (which is interpreted to be analogous to image quality assessment). Niu is relied upon for teaching a Siamese network and associated losses for finetuning a model for image quality analysis which can be mapped as the second model training to generate the third model. Wang is lastly relied upon for teaching that image quality analysis can be applied to video frames to determine video quality. The combination of the four cited art teaches the claimed subject matter of the independent claims. Regarding claim 1, Talebi teaches a method for machine learning model retraining for video quality assessment, the method comprising: retraining a first machine learning model, trained for image object detection (Page 3, Fig. 3, section 2.1: “Our image quality predictor is built on image classifier architectures [34]. We explore various classifier architectures such as VGG16 [32], Inception-v2 [33], and MobileNet [13], which are primarily used for object detection”. The initial machine model (which is the baseline image classifier network) is retrained for image quality determination using the AVA dataset “Baseline CNN weights are initialized by training on the ImageNet dataset [22], and the last fully-connected layer is initialized randomly. All NIMA weights are found by retraining on the AVA dataset” (section 2.1)), receiving a first (Page 3, section 2.1.: “Baseline CNN weights are initialized by training on the Image Net dataset [22], and the last fully-connected layer is initialized randomly. All NIMA weights are found by retraining on the AVA dataset”. The image Net dataset is composed of user generated content and would constitute the first dataset which would contain a first frame (images are functionally similar to frames of a video). The AVA dataset is made from the image collected from the dpchallenge which is a digital photography challenge, which is a type of user generated content, and would constitute the second dataset which would contain a second frame (images are functionally similar to frames of a video); retraining the (Page 3, Fig. 3, section 2.1: “Our image quality predictor is built on image classifier architectures [34]. We explore various classifier architectures such as VGG16 [32], Inception-v2 [33], and MobileNet [13], which are primarily used for object detection”. The initial machine model (which is the baseline image classifier network) is retrained for image quality determination using the AVA dataset “Baseline CNN weights are initialized by training on the ImageNet dataset [22], and the last fully-connected layer is initialized randomly. All NIMA weights are found by retraining on the AVA dataset” (section 2.1)). Talebi does not expressly disclose using two-stage transfer learning to improve the task of video quality analysis (the retraining of the first model the create a second model with a different task before finetuning is indicative of two-stage transfer learning). That is, Talebi does not expressly disclose that the first claimed retraining step is performed to produce a second machine learning model for technical content assessment using a first retraining data set or that it is this second machine learning model that is retrained to obtain the third machine learning model or that the first and second retraining data sets are different. However, Shan teaches the use two-stage transfer learning using 2 separate retraining datasets to improve an image analysis task (Fig. 1, Page 4 section 2.3a: First stage transfer learning is described where an initial pre-trained CNN (for image classification) is retrained for a separate task (saliency prediction) using a first dataset, SALICON. Then that model is finetuned with another dataset FiWI for the same task (web page saliency prediction). Section 3.1 details the datasets). Shan describes the two-stage transfer learning as such, “In fact, the two-stage transfer learning can be regarded as a task transfer in the first stage and a domain transfer in the second stage, respectively. The experimental results indicate that the proposed two-stage transfer learning of end-to-end CNN can obtain a substantial performance improvement for webpage saliency prediction”. Using this two-step method implies improved performance for task changing using transfer learning. Implementing these teachings in Talebi would result in a second model being produced by retaining the first model before it is finetuned to create the third model. The second model would have a task similar to the third model since the second finetuning simply implements a domain change as defined by Shan, and since Talebi’s third model is trained for image quality analysis, the second model would have to be trained for the same task, which would include the claimed “technical content assessment” (which is interpreted to be analogous to image quality assessment). At the time the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to modify Talebi’s transfer learning to include Shan’s two-stage transfer learning because such a modification is taught, suggested, or motivated by the art. More specifically, the motivation to modify Talebi to include Shan is expressly provided by Shan, stating that “The experimental results indicate that the proposed two-stage transfer learning of end-to-end CNN can obtain a substantial performance improvement” (Abstract). Therefore, it would have been obvious to one of ordinary skill in the art at the time of the invention to modify Talebi’s transfer learning to include Shan’s two-stage transfer learning with the motivation of substantial performance improvement. The person of ordinary skill in the art would have recognized the benefit of performance improvement. The combination of Talebi and Shan does not expressly disclose producing first output by processing the first frame using a first copy of the second machine learning model and second output by processing the second frame using a second copy of the second machine learning model; determining loss information associated with the first frame and the second frame based on the first output and the second output; and training a model based on the loss. However, Niu teaching producing first output by processing a first frame using a first copy of a machine learning model and second output by processing a second frame using a second copy of a second machine learning model (Page 101587, Fig. 3, section B: “In this paper, we propose a new SCNN whose structure is shown in Fig. 3. The SCNN consists of two parts: subnetworks I and II. Subnetwork I consists of two identical branches using five stacked convolutional structures for image feature extraction”. The Siamese network has a first and the second network branch that share weights and have similar structure effectively making the branches copies of each other. The first network has a first frame input to generate a first output, and the second network has a second frame input to generate a second output), determining loss information associated with the first frame and the second frame based on the first output and the second output (Page 101588, section B: “The network uses cross entropy as loss function, and its formula is as follows: (EQ8), where N represents the number of image patch pairs; y(i) =[y(i)2 ] is a two-dimensional vector used to indicate the quality of two images”. The outputs of the pairs are used to calculate a loss), and training a model based on the loss (Page 101588, section B: “SCNN is iteratively trained over multiple epochs, and an epoch is defined as traversing the entire training set. In each epoch, the training set is divided into multiple mini-batches for batch optimization”. Loss functions are used to train the network; this section follows description of the loss function implying this loss function is used for the multiple epochs). At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify the combination of Talebi and Shan’s image quality analysis system to include Niu’s Siamese network and training method for image quality analysis because such a modification is taught, suggested, or motivated by the art. More specifically, the motivation to modify the combination of Talebi and Shan to include Niu is expressly provided by Niu, stating that for image quality analysis systems that deal with distortions, it performs better (Page 101590, 1st column: “For the single-distortion-type experiments on the TID2013 Database, the proposed LRSN model achieves improved performance compared to the 6 comparison no-reference IQA models, and the improvement in WN distortion type is significant. For WN distortion type, the performance of the proposed LRSN model is close to the best performance among the four full-reference IQA metrics”). Therefore, it would have been obvious to one of ordinary skill in the art at the time of the invention to modify the combination of Talebi and Shan’s image quality analysis system to include Niu’s Siamese training method for image quality analysis with the motivation of improving image quality analysis performance. The person of ordinary skill in the art would have recognized the benefit of improved image quality analysis performance. The combination of Talebi, Shan, and Niu does not expressly disclose the images coming from video data. However, Wang teaches that images can be derived from video content for determining video quality ([0012]: “Images can be derived from video”. [0023]: “Finally, the image quality score is predicted by a loss function”). At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify Talebi’s data acquisition to include Wang’s acquisition of images from videos because such a modification is the result of applying a known technique to a known device ready for improvement to yield predictable results. More specifically, Wang’s acquisition of images from videos permits images to be acquired from video for quality determination. This known benefit in Wang is applicable to Talebi’s data acquisition as they both share characteristics and capabilities, namely, they are directed to quality determination. Therefore, it would have been recognized that modifying Talebi’s data acquisition to include Wang’s acquisition of images from videos would have yielded predictable results because (i) the level of ordinary skill in the art demonstrated by the references applied shows the ability to incorporate Wang’s acquisition of images from videos in quality determination and (ii) the benefits of such a combination would have been recognized by those of ordinary skill in the art. Regarding claim 8, the combination of Talebi, Shan, Niu, and Wang teaches the method of claim 1, in addition, Wang further teaches wherein the first frame and the second frame are received as a frame pair in which the first frame is at a first quality level and the second frame is at a second quality level (Fig. 2: Two images, a first image with a first quality (reference images) and a second image with a second quality (distorted images) can be seen being input into a network to train the network to determine image quality ([0023]: “Finally, the image quality score is predicted by a loss function”. Images are seen as functionally similar to video frames). At the time the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to substitute Talebi’s training method with Wang’s training method because such a modification is the result of simple substitution of one known element for another producing a predictable result. More specifically, Talebi’s training method and Wang’s training method perform the same general and predictable function, the predictable function being training a neural network to determine the quality of an image. Since each individual element and its function are shown in the prior art, albeit shown in separate references, the difference between the claimed subject matter and the prior art rests not on any individual element or function but in the very combination itself - that is in the substitution of Talebi’s training method by replacing it with Wang’s training method. Thus, the simple substitution of one known element for another producing a predictable result renders the claim obvious. Regarding claim 9, the combination of Talebi, Shan, Niu, and Wang teaches the method of claim 8, in addition, Talebi further teaches wherein receiving the first retraining data set and the second retraining data set comprises: extracting frames from each of the first user generated video content and the second user generated video content, wherein the frames include the first frame and the second frame (Page 3, section 2.1.: “Baseline CNN weights are initialized by training on the Image Net dataset [22], and the last fully-connected layer is initialized randomly. All NIMA weights are found by retraining on the AVA dataset”. The image Net dataset is composed of user generated content and would constitute the first dataset which would contain a first frame (images are functionally similar to frames of a video). The AVA dataset is made from the image collected from the dpchallenge which is a digital photography challenge, which is a type of user generated content, and would constitute the second dataset which would contain a second frame (images are functionally similar to frames of a video). Regarding claim 10, the combination of Talebi, Shan, and Niu teaches the method of claim 1, in addition, Talebi further teaches wherein the first user generated content is the second user generated content (Page 3, section 2.1.: “Baseline CNN weights are initialized by training on the Image Net dataset [22], and the last fully-connected layer is initialized randomly. All NIMA weights are found by retraining on the AVA dataset”. The image Net dataset is composed of user generated content and would constitute the first dataset which would contain a first frame (images are functionally similar to frames of a video). The AVA dataset is made from the image collected from the dpchallenge which is a digital photography challenge, which is a type of user generated content, and would constitute the second dataset which would contain a second frame (images are functionally similar to frames of a video. The first user generated content is the second user generated content since they are both the same type of data (user generated images)). Regarding claim 11, the content of claim 11 is similar to the content of claim 1, therefore it is rejected for the same reasons of obviousness as claim 1. Regarding claim 12, the content of claim 12 is similar to the content of claim 1, therefore it is rejected for the same reasons of obviousness as claim 1. Regarding claim 13, the combination of Talebi, Shan, Niu, and Wang teaches the method of claim 11, in addition, Niu further teaches wherein determining the loss information associated with the first frame and the second frame based on the first output and the second output comprises: determining the loss information based on one or more of (Page 101588, section B: “The network uses cross entropy as loss function, and its formula is as follows: (EQ8), where N represents the number of image patch pairs; y(i) =[y(i)2 ] is a two-dimensional vector used to indicate the quality of two images”. The outputs of the pairs are used to calculate a loss. The loss is based on the model components since the output of the image pairs are processed using the model and the loss is calculated using the outputs. The “or” language means the list is in the alternative and only 1 of the listed items need be taught for a case of obviousness). The rationale for this combination is similar to the rationale for the claim1 combination with Niu due to similar methods of combination (using a Siamese network and loss functions associated with it for image quality) and similar benefits (improved image analysis). Regarding claim 14, the content of claim 14 is similar to the content of claim 9, therefore it is rejected for the same reasons of obviousness as claim 9. Regarding claim 16, the content of claim 16 is similar to the content of claim 1, therefore it is rejected for the same reasons of obviousness as claim 1. Regarding claim 20, the content of claim 20 is similar to the content of claim 8, therefore it is rejected for the same reasons of obviousness as claim 8. Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Talebi et al. (“Learned Perceptual Image Enhancement” Hereinafter “Talebi”) in view of Shan et al. (“Two-Stage Transfer Learning of End-to-End Convolutional Neural Networks for Webpage Saliency Prediction” Hereinafter “Shan”) in further view of Niu et al. (“Siamese-Network-Based Learning to Rank for No-Reference 2D and 3D Image Quality Assessment” Hereinafter “Niu”) in further view of Wang et al. (US 20190122115 A1 Hereinafter “Wang”) in further view of Dong et al. (“Triplet Loss in Siamese Network for Object Tracking “Hereinafter “Dong”). Regarding claim 4, the combination of Talebi, Shan, Niu, and Wang teaches the method of claim 1, in addition, Niu further teaches wherein the loss information (Page 101588, section B: “The network uses cross entropy as loss function, and its formula is as follows: (EQ8), where N represents the number of image patch pairs; y(i) =[y(i)2 ] is a two-dimensional vector used to indicate the quality of two images”. The outputs of the pairs are used to calculate a loss). At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify the combination of Talebi and Shan’s image quality analysis system to include Niu’s Siamese network and training method for image quality analysis because such a modification is taught, suggested, or motivated by the art. More specifically, the motivation to modify the combination of Talebi and Shan to include Niu is expressly provided by Niu, stating that for image quality analysis systems that deal with distortions, it performs better (Page 101590, 1st column: “For the single-distortion-type experiments on the TID2013 Database, the proposed LRSN model achieves improved performance compared to the 6 comparison no-reference IQA models, and the improvement in WN distortion type is significant. For WN distortion type, the performance of the proposed LRSN model is close to the best performance among the four full-reference IQA metrics”). Therefore, it would have been obvious to one of ordinary skill in the art at the time of the invention to modify the combination of Talebi and Shan’s image quality analysis system to include Niu’s Siamese training method for image quality analysis with the motivation of improving image quality analysis performance. The person of ordinary skill in the art would have recognized the benefit of improved image quality analysis performance. The combination of Talebi, Shan, Niu, and Wang does not expressly disclose training the machine learning model using the rank loss information and using cross entropy loss information. However, Dong teaches disclose training the machine learning model using the rank loss information and using cross entropy loss information (Page 3, Fig. 1: The figure shows a Siamese network that uses 2 types off loss, logistical (functionally similar to cross-entropy) and Triplet loss (functionally similar to rank loss). At the time the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to substitute the combination of Talebi, Shan, Niu, and Wang’s training method loss with Dong’s use of two losses because such a modification is the result of simple substitution of one known element for another producing a predictable result. More specifically, the combination of Talebi, Shan, Niu, and Wang’s training method loss and Dong’s use of two losses perform the same general and predictable function, the predictable function being training a neural network using losses to determine the quality of an image. Since each individual element and its function are shown in the prior art, albeit shown in separate references, the difference between the claimed subject matter and the prior art rests not on any individual element or function but in the very combination itself - that is in the substitution of the combination of Talebi, Shan, Niu, and Wang’s training method loss by replacing it with and Dong’s use of two losses. Thus, the simple substitution of one known element for another producing a predictable result renders the claim obvious. Claims 5 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Talebi et al. (“Learned Perceptual Image Enhancement” Hereinafter “Talebi”) in view of Shan et al. (“Two-Stage Transfer Learning of End-to-End Convolutional Neural Networks for Webpage Saliency Prediction” Hereinafter “Shan”) in further view of Niu et al. (“Siamese-Network-Based Learning to Rank for No-Reference 2D and 3D Image Quality Assessment” Hereinafter “Niu”) in further view of (“Video Quality Adaptation for Limiting Transcoding Energy Consumption in Video Servers “Hereinafter “Lee”). Regarding claim 5, the combination of Talebi, Shan, Niu, and Wang teaches the method of claim 1, comprising: The combination of Talebi, Shan, Niu, and Wang does not expressly disclose transcoding an input video stream of user generated video content using the third machine learning model. However, Lee teaches transcoding an input video stream using a machine learning model (Pages 126255, section 3: “To address this, we develop a transcoding parameter selection (TPS) algorithm to determine which versions are transcoded in transcoding session by considering video popularity and quality” (Emphasis added). “TPS algorithms take an energy limit parameter (Elimit) and video segment parameters (Ctotali, j and Gi, j) to determine transcodable versions for each segment. When new video clips are released or uploaded so that workloads in transcoding session change, TPS is executed to find the best transcoding parameters for the segments in a current job queue” (Emphasis added) (126258-126259). The TSP uses video segment parameters C and G to find the best transcoding parameters, G is defined on page 126257 as taking quality of a video into account, so different qualities of videos would result in different transcoding parameters, for transcoding the video “From this table, we observe that the required transcoding time is quite different depending on the resolution of the source version”. Including the citation before, the parameters are selected and if they observed the transcoding time they transcoded the input video (Page 126259 section 5)). At the time the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to substitute the Lee’s video quality model with the combination of Talebi, Shan, Niu, and Wang’s video quality model because such a modification is the result of simple substitution of one known element for another producing a predictable result. More specifically, Lee’s video quality model and with the combination of Talebi, Shan, Niu, and Wang’s video quality model perform the same general and predictable function, the predictable function being using a trained model to predict video quality to affect transcoding parameters. Since each individual element and its function are shown in the prior art, albeit shown in separate references, the difference between the claimed subject matter and the prior art rests not on any individual element or function but in the very combination itself - that is in the substitution of Lee’s video quality model by replacing it with the combination of Talebi, Shan, Niu, and Wang’s video quality model. Thus, the simple substitution of one known element for another producing a predictable result renders the claim obvious Regarding claim 18, the combination of Talebi, Shan, Niu, and Wang teaches the method of claim 16, The combination of Talebi, Shan, Niu, and Wang does not expressly disclose wherein an input video stream of user generated video content is transcoded using the third machine learning model. However, Lee teaches transcoding an input video stream using a machine learning model (Pages 126255, section 3: “To address this, we develop a transcoding parameter selection (TPS) algorithm to determine which versions are transcoded in transcoding session by considering video popularity and quality” (Emphasis added). “TPS algorithms take an energy limit parameter (Elimit) and video segment parameters (Ctotali, j and Gi,j ) to determine transcodable versions for each segment. When new video clips are released or uploaded so that workloads in transcoding session change, TPS is executed to find the best transcoding parameters for the segments in a current job queue” (Emphasis added) (126258-126259). The TSP uses video segment parameters C and G to find the best transcoding parameters, G is defined on page 126257 as taking quality of a video into account, so different qualities of videos would result in different transcoding parameters, for transcoding the video “From this table, we observe that the required transcoding time is quite different depending on the resolution of the source version”. Including the citation before, the parameters are selected and if they observed the transcoding time they transcoded the input video (Page 126259 section 5) and a multi-stage transcoding pipeline, wherein a first stage of the multi-stage transcoding pipeline completes processing of the input video stream before a second stage of the multi-stage transcoding pipeline completes processing of the input video stream (Fig. 1, page 126256: “1) A TPS algorithm is executed to determine which versions can be transcoded in transcoding session for each video segment. 2) Each video then goes through a transcoding session to transcode the versions determined to be transcoded by the TPS algorithm”. The first stage must finish before the second stage due to the second stage needing parameters from the first stage). At the time the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to substitute the Lee’s video quality model with the combination of Talebi, Shan, Niu, and Wang’s video quality model because such a modification is the result of simple substitution of one known element for another producing a predictable result. More specifically, Lee’s video quality model and with the combination of Talebi, Shan, Niu, and Wang’s video quality model perform the same general and predictable function, the predictable function being using a trained model to predict video quality to affect transcoding parameters. Since each individual element and its function are shown in the prior art, albeit shown in separate references, the difference between the claimed subject matter and the prior art rests not on any individual element or function but in the very combination itself - that is in the substitution of Lee’s video quality model by replacing it with the combination of Talebi, Shan, Niu, and Wang’s video quality model. Thus, the simple substitution of one known element for another producing a predictable result renders the claim obvious. Allowable Subject Matter Claims 2-3, 6-7, 15, 17, and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: SETHURAMAN et al. (US 20190075299 A1) teaches a Siamese network for a transcoding pipeline Schmidt et al. (US 20150381690 A1) teaches transcoding videos from a mezzanine format based on adaptive parameters. Lin et al. (US 9813706 B1) teaches transcoding using mezzanine format Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEFANO A DARDANO whose telephone number is (703)756-4543. The examiner can normally be reached Monday - Friday 11:00 - 7:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Greg Morse can be reached at (571) 272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /STEFANO ANTHONY DARDANO/ Examiner, Art Unit 2663 /GREGORY A MORSE/Supervisory Patent Examiner, Art Unit 2698
Read full office action

Prosecution Timeline

Feb 11, 2025
Application Filed
Sep 17, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737888
DATA PROCESSING METHOD AND DEVICE
2y 4m to grant Granted Sep 15, 2026
Patent 12725269
CELL DETECTION METHOD AND APPARATUS, DEVICE, READABLE STORAGE MEDIUM, AND PROGRAM PRODUCT
2y 10m to grant Granted Sep 01, 2026
Patent 12718402
DEVICE AND METHOD FOR TRAINING A MACHINE LEARNING MODEL FOR GENERATING DESCRIPTOR IMAGES FOR IMAGES OF OBJECTS
3y 9m to grant Granted Aug 25, 2026
Patent 12705757
Image Processor
2y 10m to grant Granted Aug 11, 2026
Patent 12688567
METHOD FOR DETECTING AND LOCALIZING A FALSIFIED AREA IN JPEG IMAGES
3y 3m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+31.9%)
2y 11m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 93 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month