Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED OFFICE ACTION
Status of Claims
Claims 1-20 are pending in this Office Action.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b) (2) (C) for any potential 35 U.S.C. 102(a) (2) prior art against the later invention.
1. Claims 1,2,7,8,9,14,15 and 20 are rejected under 35 U.S.C 103 as being patentable over Suri et al. (USPUB 20160034786) in view of Ning Yu et al.( NPL Doc: “Learning to Detect Multiple Photographic Defects,” 8th March 2018, 2018 IEEE Winter Conference on Applications of Computer Vision, Pages 1387-1394).
As per claim 1, Suri et al. teaches A method comprising: training a binary model ( Paragraph [0059]- “…Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression,…” ) at least in part on a dataset of images labeled to indicate exposure within the images ( Paragraphs [0081-0082]- “… extracting module 118 extracts low level and high level features from the video data for training models, as described above. Feature extraction may describe the process of identifying different attributes of video data and leveraging those features for additional processing. Features may be extracted per video frame, video segment, video file, and/or video collection level. The features may include exposure quality, saturation quality, hue variety,…”) , such that the binary model, when trained, classifies an image based on whether the image includes an exposure defect ( Paragraph [0044]- “…technical quality and subjective importance, as perceived from a human labeling the video data. For example, on a five point scale, the combination of strong technical quality (e.g., full exposure, even color distribution, minimal camera motion, bright colors, focused faces, etc.) and subject importance above a predetermined threshold (e.g., main character, clear audio, minimal camera motion and/or object motion, etc.) may result in desirable video data, or a desirability score closer to five. In contrast, poor technical quality (e.g., poor exposure, uneven color distribution, significant camera motion, dark picture, etc.) and/or subjective importance below a predetermined threshold (e.g., unimportant character, muffled audio, significant camera motion and/or object motion, etc.) may result in undesirable video data, or a desirability score closer to zero. In at least some examples, video data may have a neutral level of desirability (e.g., the video data is not desirable or undesirable), or a desirability score near 2. The receiving module 202 may provide the received video data to the extracting module 118 for feature extraction before training the classifier and scoring model in the learning module 204….”) ; and storing the binary model and the classification model for identifying and classifying exposure defects ( Paragraphs [0044-0045]“…the combination of strong technical quality (e.g., full exposure, even color distribution, minimal camera motion, bright colors, focused faces, etc.) and subject importance above a predetermined threshold (e.g., main character, clear audio, minimal camera motion and/or object motion, etc.) may result in desirable video data, or a desirability score closer to five. In contrast, poor technical quality (e.g., poor exposure, uneven color distribution, significant camera motion, dark picture, etc.) and/or subjective importance below a predetermined threshold (e.g., unimportant character, muffled audio, significant camera motion and/or object motion, etc.) may result in undesirable video data, or a desirability score closer to zero. In at least some examples, video data may have a neutral level of desirability (e.g., the video data is not desirable or undesirable), or a desirability score near 2. The receiving module 202 may provide the received video data to the extracting module 118 for feature extraction before training the classifier and scoring model in the learning module 204….”) AND Paragraphs [0059-0060]- “…The learning module 204 may be configured to train a classifier and a scoring model based on the low level, high level, and derivative feature values. The classifier includes a plurality of classifiers that may be used to generate a plurality of high level semantic feature values. The classifier may be used to estimate a probability that a video frame, video segment, video file, and/or video collection belongs to at least one of the categories in the predefined set of semantic categories (e.g., indoor, outdoor, mountain, lake, city, country, home, party, sporting event, zoo, concert, etc.). The classifier may be trained by applying models such as Linear Support Vector Machine (SVM), Fast Tree Binary Classification, Fast Rank Binary Classification, …”) .
Suri et al. does not explicitly teach training a classification model at least in part on a dataset of images having exposure defects labeled to indicate exposure scores or exposure defect classifications, such that the classification model, when trained, classifies the image based on a level of exposure;
However, within analogous art, Ning Yu et al. teaches training a classification model at least in part on a dataset of images having exposure defects labeled to indicate exposure scores or exposure defect classifications ( Page 1391- Col.1- “….to train a deep convolutional neural network (CNN) to directly learn high-level understanding of photographic defects from human judgments…. seven defects at the same time. These defects are related to low-level photo properties such as color, exposure, noise, and blur, and high-level properties such as faces, humans, and compositional balance. We note that both low- and high-level content features may be useful. Therefore, we use a multicolumn CNN, in which the earlier layers of the network are shared across all the defects to learn defect-agnostic features, and in later layers, a separate branch is dedicated to each defect to capture defect-specific information….”) , such that the classification model, when trained, classifies the image based on a level of exposure ( Page 1393- Col. 2- “…Although our model was trained on our dataset of defective images in the wild, we can also validate our trained model on an easier dataset of synthetically generated global defects. We separately generate defective images for under exposure, over exposure, over/under saturation, Gaussian… noise, and spatially invariant motion blur. We first select for each defect all of the defect-free testing images (there are between 420 and 940 such images)…” AND Table. 4 & 5) ;
One of ordinary skill in the art would have been motivated to combine the teaching of Ning Yu et al. within the modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. because the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. provides a system and method for implementing image exposure defect detection with neural network models.
Therefore, it would have been obvious for one in the ordinary skills in the art before the effective filing date of the claimed invention to implement the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. within the Computerized machine learning of interesting video sections mentioned by Suri et al. for implementation of a system and method for image exposure defect detection with neural network models.
As per claim 2,Combination of Suri et al. and Ning Yu et al. teach claim 1,
Suri et al. teaches wherein: the binary model is a first neural network (FIG. 2- TRAINING MODULE ( 116) AND Paragraph [0059]- “…Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression, Averaged Perception, etc., to the feature values extracted or derived, as described above….”) ; and the classification model is a second neural network ( FIG. 2- CLASSIFYING MODULE ( 124) AND Paragraph [0061]- “… the feature values resulting from applying the classifier to the feature values allows the scoring model to account for variations in categories of video data. For example, by using the feature values resulting from applying the classifier,…”) .
As per claim 7, Combination of Suri et al. and Ning Yu et al. teach claim 1,
Suri et al. teaches wherein the binary model is a feature-based model( Paragraph [0059]- “…The classifier may be trained by applying models such as Linear Support Vector Machine (SVM), Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression, Averaged Perception, etc., to the feature values extracted or derived, as described above. The learned classifier may be stored in the classifying module 124….”) .
As per claim 8, Suri et al. teaches One or more computer storage media storing computer-useable instructions that, when executed by one or more computing devices ( Paragraph [0067]- “… In the context of software, the blocks represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations….”) , cause the one or more computing devices to perform operations ( Paragraph [0067]- “…computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular abstract data types. …”) comprising:
training a binary model ( Paragraph [0059]- “…Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression,…” ) at least in part on a dataset of images labeled to indicate exposure within the images( Paragraphs [0081-0082]- “… extracting module 118 extracts low level and high level features from the video data for training models, as described above. Feature extraction may describe the process of identifying different attributes of video data and leveraging those features for additional processing. Features may be extracted per video frame, video segment, video file, and/or video collection level. The features may include exposure quality, saturation quality, hue variety,…”), such that the binary model, when trained, classifies an image based on whether the image includes an exposure defect( Paragraph [0044]- “…technical quality and subjective importance, as perceived from a human labeling the video data. For example, on a five point scale, the combination of strong technical quality (e.g., full exposure, even color distribution, minimal camera motion, bright colors, focused faces, etc.) and subject importance above a predetermined threshold (e.g., main character, clear audio, minimal camera motion and/or object motion, etc.) may result in desirable video data, or a desirability score closer to five. In contrast, poor technical quality (e.g., poor exposure, uneven color distribution, significant camera motion, dark picture, etc.) and/or subjective importance below a predetermined threshold (e.g., unimportant character, muffled audio, significant camera motion and/or object motion, etc.) may result in undesirable video data, or a desirability score closer to zero. In at least some examples, video data may have a neutral level of desirability (e.g., the video data is not desirable or undesirable), or a desirability score near 2. The receiving module 202 may provide the received video data to the extracting module 118 for feature extraction before training the classifier and scoring model in the learning module 204….”); and storing the binary model and the at least two classification models for identifying and classifying exposure defects( Paragraphs [0044-0045]“…the combination of strong technical quality (e.g., full exposure, even color distribution, minimal camera motion, bright colors, focused faces, etc.) and subject importance above a predetermined threshold (e.g., main character, clear audio, minimal camera motion and/or object motion, etc.) may result in desirable video data, or a desirability score closer to five. In contrast, poor technical quality (e.g., poor exposure, uneven color distribution, significant camera motion, dark picture, etc.) and/or subjective importance below a predetermined threshold (e.g., unimportant character, muffled audio, significant camera motion and/or object motion, etc.) may result in undesirable video data, or a desirability score closer to zero. In at least some examples, video data may have a neutral level of desirability (e.g., the video data is not desirable or undesirable), or a desirability score near 2. The receiving module 202 may provide the received video data to the extracting module 118 for feature extraction before training the classifier and scoring model in the learning module 204….”) AND Paragraphs [0059-0060]- “…The learning module 204 may be configured to train a classifier and a scoring model based on the low level, high level, and derivative feature values. The classifier includes a plurality of classifiers that may be used to generate a plurality of high level semantic feature values. The classifier may be used to estimate a probability that a video frame, video segment, video file, and/or video collection belongs to at least one of the categories in the predefined set of semantic categories (e.g., indoor, outdoor, mountain, lake, city, country, home, party, sporting event, zoo, concert, etc.). The classifier may be trained by applying models such as Linear Support Vector Machine (SVM), Fast Tree Binary Classification, Fast Rank Binary Classification, …”).
Suri et al. does not explicitly teach training at least two classification models at least in part on a dataset of images having exposure defects labeled to indicate exposure scores or exposure defect classifications, such that a first classification model, when trained, is configured to classify the image based on a level of underexposure, and such that a second classification model, when trained, is configured to classify the image based on a level of overexposure;
However, within analogous art, Ning Yu et al. teaches training at least two classification models at least in part on a dataset of images having exposure defects labeled to indicate exposure scores or exposure defect classifications ( Page 1391- Col.1- “….to train a deep convolutional neural network (CNN) to directly learn high-level understanding of photographic defects from human judgments…. seven defects at the same time. These defects are related to low-level photo properties such as color, exposure, noise, and blur, and high-level properties such as faces, humans, and compositional balance. We note that both low- and high-level content features may be useful. Therefore, we use a multicolumn CNN, in which the earlier layers of the network are shared across all the defects to learn defect-agnostic features, and in later layers, a separate branch is dedicated to each defect to capture defect-specific information….”) , such that a first classification model, when trained, is configured to classify the image based on a level of underexposure ( Page 1393- Col. 2- “…Although our model was trained on our dataset of defective images in the wild, we can also validate our trained model on an easier dataset of synthetically generated global defects. We separately generate defective images for under exposure, over exposure, over/under saturation, Gaussian… noise, and spatially invariant motion blur. We first select for each defect all of the defect-free testing images (there are between 420 and 940 such images)…” AND Table. 4 & 5) ; and such that a second classification model, when trained, is configured to classify the image based on a level of overexposure ( Pages 1293-1394- Col. 2-Col. 1- “…We separately generate defective images for under exposure, over exposure, over/under saturation, Gaussian noise, and spatially invariant motion blur. We first select for each defect all of the defect-free testing images (there are between 420 and 940 such images). For each such image, we synthesize a sequence of defective images with either11 or 21 different levels of a global parameter, where the number of levels is chosen to be consistent with the class structure in our user dataset discussed in Section 3.1….”) ;
One of ordinary skill in the art would have been motivated to combine the teaching of Ning Yu et al. within the modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. because the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. provides a system and method for implementing image exposure defect detection with neural network models.
Therefore, it would have been obvious for one in the ordinary skills in the art before the effective filing date of the claimed invention to implement the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. within the Computerized machine learning of interesting video sections mentioned by Suri et al. for implementation of a system and method for image exposure defect detection with neural network models.
As per claim 9, Combination of Suri et al. and Ning Yu et al. teach claim 8,
Suri et al. teaches wherein: the binary model is a first neural network (FIG. 2- TRAINING MODULE ( 116) AND Paragraph [0059]- “…Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression, Averaged Perception, etc., to the feature values extracted or derived, as described above….”) ; and the at least two classification models comprise a second neural network ( FIG. 2- CLASSIFYING MODULE ( 124) AND Paragraph [0061]- “… the feature values resulting from applying the classifier to the feature values allows the scoring model to account for variations in categories of video data. For example, by using the feature values resulting from applying the classifier,…”) .
As per claim 14, Combination of Suri et al. and Ning Yu et al. teach claim 8,
Suri et al. teaches wherein each of the binary model and the at least two classification models are feature-based models ( Paragraph [0059]- “…The classifier may be trained by applying models such as Linear Support Vector Machine (SVM), Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression, Averaged Perception, etc., to the feature values extracted or derived, as described above. The learned classifier may be stored in the classifying module 124….”) .
As per claim 15, Suri et al. teaches A system comprising one or more processors and memory configured to provide computer program instructions to the one or more processors( Paragraph [0067]- “… In the context of software, the blocks represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations….”), the computer program instructions causing the one or more processors( Paragraph [0067]- “…computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular abstract data types. …”) to: train a first neural network binary model ( Paragraph [0059]- “…Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression,…” ) at least in part on a dataset of images labeled to indicate exposure within the images ( Paragraphs [0081-0082]- “… extracting module 118 extracts low level and high level features from the video data for training models, as described above. Feature extraction may describe the process of identifying different attributes of video data and leveraging those features for additional processing. Features may be extracted per video frame, video segment, video file, and/or video collection level. The features may include exposure quality, saturation quality, hue variety,…”), such that the first neural network binary model, when trained, classifies an image based on whether the image includes an exposure defect ( Paragraph [0044]- “…technical quality and subjective importance, as perceived from a human labeling the video data. For example, on a five point scale, the combination of strong technical quality (e.g., full exposure, even color distribution, minimal camera motion, bright colors, focused faces, etc.) and subject importance above a predetermined threshold (e.g., main character, clear audio, minimal camera motion and/or object motion, etc.) may result in desirable video data, or a desirability score closer to five. In contrast, poor technical quality (e.g., poor exposure, uneven color distribution, significant camera motion, dark picture, etc.) and/or subjective importance below a predetermined threshold (e.g., unimportant character, muffled audio, significant camera motion and/or object motion, etc.) may result in undesirable video data, or a desirability score closer to zero. In at least some examples, video data may have a neutral level of desirability (e.g., the video data is not desirable or undesirable), or a desirability score near 2. The receiving module 202 may provide the received video data to the extracting module 118 for feature extraction before training the classifier and scoring model in the learning module 204….”); and store the first neural network binary model and the second neural network classification model for identifying and classifying exposure defects ( Paragraphs [0044-0045]“…the combination of strong technical quality (e.g., full exposure, even color distribution, minimal camera motion, bright colors, focused faces, etc.) and subject importance above a predetermined threshold (e.g., main character, clear audio, minimal camera motion and/or object motion, etc.) may result in desirable video data, or a desirability score closer to five. In contrast, poor technical quality (e.g., poor exposure, uneven color distribution, significant camera motion, dark picture, etc.) and/or subjective importance below a predetermined threshold (e.g., unimportant character, muffled audio, significant camera motion and/or object motion, etc.) may result in undesirable video data, or a desirability score closer to zero. In at least some examples, video data may have a neutral level of desirability (e.g., the video data is not desirable or undesirable), or a desirability score near 2. The receiving module 202 may provide the received video data to the extracting module 118 for feature extraction before training the classifier and scoring model in the learning module 204….”) AND Paragraphs [0059-0060]- “…The learning module 204 may be configured to train a classifier and a scoring model based on the low level, high level, and derivative feature values. The classifier includes a plurality of classifiers that may be used to generate a plurality of high level semantic feature values. The classifier may be used to estimate a probability that a video frame, video segment, video file, and/or video collection belongs to at least one of the categories in the predefined set of semantic categories (e.g., indoor, outdoor, mountain, lake, city, country, home, party, sporting event, zoo, concert, etc.). The classifier may be trained by applying models such as Linear Support Vector Machine (SVM), Fast Tree Binary Classification, Fast Rank Binary Classification, …”) .
Suri et al. does not explicitly teach train a second neural network classification model at least in part on a dataset of images having exposure defects labeled to indicate exposure scores or exposure defect classifications, such that the second neural network classification model, when trained, classifies the image from the first neural network binary model based on a level of exposure;
However, within analogous art, Ning Yu et al. teaches train a second neural network classification model at least in part on a dataset of images having exposure defects labeled to indicate exposure scores or exposure defect classifications ( Page 1391- Col.1- “….to train a deep convolutional neural network (CNN) to directly learn high-level understanding of photographic defects from human judgments…. seven defects at the same time. These defects are related to low-level photo properties such as color, exposure, noise, and blur, and high-level properties such as faces, humans, and compositional balance. We note that both low- and high-level content features may be useful. Therefore, we use a multicolumn CNN, in which the earlier layers of the network are shared across all the defects to learn defect-agnostic features, and in later layers, a separate branch is dedicated to each defect to capture defect-specific information….”) , such that the second neural network classification model, when trained, classifies the image from the first neural network binary model based on a level of exposure ( Page 1393- Col. 2- “…Although our model was trained on our dataset of defective images in the wild, we can also validate our trained model on an easier dataset of synthetically generated global defects. We separately generate defective images for under exposure, over exposure, over/under saturation, Gaussian… noise, and spatially invariant motion blur. We first select for each defect all of the defect-free testing images (there are between 420 and 940 such images)…” AND Table. 4 & 5) ;
One of ordinary skill in the art would have been motivated to combine the teaching of Ning Yu et al. within the modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. because the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. provides a system and method for implementing image exposure defect detection with neural network models.
Therefore, it would have been obvious for one in the ordinary skills in the art before the effective filing date of the claimed invention to implement the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. within the Computerized machine learning of interesting video sections mentioned by Suri et al. for implementation of a system and method for image exposure defect detection with neural network models.
As per claim 20, Combination of Suri et al. and Ning Yu et al. teach claim 15,
Suri et al. teaches wherein at least one of the first neural network binary model and the second neural network classification model is a feature-based model ( Paragraph [0059]- “…The classifier may be trained by applying models such as Linear Support Vector Machine (SVM), Fast Tree Binary Classification, Fast Rank Binary Classification, Logistic Regression, Averaged Perception, etc., to the feature values extracted or derived, as described above. The learned classifier may be stored in the classifying module 124….”) .
2. Claims 3 and 10 are rejected under 35 U.S.C 103 as being patentable over Suri et al. (USPUB 20160034786) in view of Ning Yu et al.( NPL Doc: “Learning to Detect Multiple Photographic Defects,” 8th March 2018, 2018 IEEE Winter Conference on Applications of Computer Vision, Pages 1387-1394) in further view of Singh et al. (USPUB 20210124993).
As per claim 3, Combination of Suri et al. and Ning Yu et al. teach claim 2,
Combination of Suri et al. and Ning Yu et al. does not explicitly teach wherein: the first neural network and the second neural network share at least one layer ; and the first neural network and the second neural network each have a separate bottom layer.
Within analogous art, Singh et al. teaches wherein: the first neural network and the second neural network share at least one layer ( Fig. 2 AND Paragraph [0041]- “…digital image classification system 102 can train one or more neural networks such as a base neural network and a classification neural network to classify digital images in few-shot tasks…” AND Paragraph [0084]) ; and the first neural network and the second neural network each have a separate bottom layer ( Paragraph [0030-0031]- “ the term “feature extractor” refers to one or more layers of a neural network that extract features relating to digital images. For example, a feature extractor can include a particular number of layers (e.g., 4 layers) including one or more fully connected and/or partially connected layers of neurons that identify and represent visible and/or unobservable characteristics of a digital image…”) .
One of ordinary skill in the art would have been motivated to combine the teaching of Singh et al. within the combined modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. and the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. because the Classifying digital images in few-shot tasks based on neural networks trained using manifold mixup regularization and self-supervision mentioned by Singh et al. provides a system and method for implementing training of classification neural network for image processing .
Therefore, it would have been obvious for one in the ordinary skills in the art before the effective filing date of the claimed invention to implement the Classifying digital images in few-shot tasks based on neural networks trained using manifold mixup regularization and self-supervision mentioned by Singh et al. within the combined modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. and the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. provides a system and method for implementing training of classification neural network for image processing .
As per claim 10, Combination of Suri et al. and Ning Yu et al. teach claim 9,
Combination of Suri et al. and Ning Yu et al. does not explicitly teach wherein: the first neural network and the second neural network share a top layer; and the first neural network and the second neural network each have a separate bottom layer.
Within analogous art, Singh et al. teaches wherein: the first neural network and the second neural network share a top layer ( Fig. 2 AND Paragraph [0041]- “…digital image classification system 102 can train one or more neural networks such as a base neural network and a classification neural network to classify digital images in few-shot tasks…” AND Paragraph [0084]); and the first neural network and the second neural network each have a separate bottom layer( Paragraph [0030-0031]- “ the term “feature extractor” refers to one or more layers of a neural network that extract features relating to digital images. For example, a feature extractor can include a particular number of layers (e.g., 4 layers) including one or more fully connected and/or partially connected layers of neurons that identify and represent visible and/or unobservable characteristics of a digital image…”) .
One of ordinary skill in the art would have been motivated to combine the teaching of Singh et al. within the combined modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. and the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. because the Classifying digital images in few-shot tasks based on neural networks trained using manifold mixup regularization and self-supervision mentioned by Singh et al. provides a system and method for implementing training of classification neural network for image processing .
Therefore, it would have been obvious for one in the ordinary skills in the art before the effective filing date of the claimed invention to implement the Classifying digital images in few-shot tasks based on neural networks trained using manifold mixup regularization and self-supervision mentioned by Singh et al. within the combined modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. and the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. provides a system and method for implementing training of classification neural network for image processing .
3. Claim 13 is rejected under 35 U.S.C 103 as being patentable over Suri et al. (USPUB 20160034786) in view of Ning Yu et al.( NPL Doc: “Learning to Detect Multiple Photographic Defects,” 8th March 2018, 2018 IEEE Winter Conference on Applications of Computer Vision, Pages 1387-1394) in further view of CHOE et al. (USPUB 20200026282).
As per claim 13, Combination of Suri et al. and Ning Yu et al. teach claim 9,
Combination of Suri et al. and Ning Yu et al. does not explicitly teach wherein the first neural network is trained using a weak supervised learning algorithm.
Within analogous art, CHOE et al. teaches wherein the first neural network is trained using a weak supervised learning algorithm ( Paragraph [0059]- “… the neural network model can include a segmentation algorithm to generate a binary image for the lane lines, e.g., by segmenting the heat map. The segmentation algorithm can include a threshold-based segmentation, a region-based segmentation, an edge detection segmentation, a segmentation based on clustering, a segmentation based on weakly-supervised learning in convolutional neural network algorithm, …”) .
One of ordinary skill in the art would have been motivated to combine the teaching of CHOE et al. within the combined modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. and the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. because the Lane/object detection and tracking perception system for autonomous vehicles mentioned by CHOE et al. provides a system and method for implementing machine learning algorithm for image processing.
Therefore, it would have been obvious for one in the ordinary skills in the art before the effective filing date of the claimed invention to implement the Lane/object detection and tracking perception system for autonomous vehicles mentioned by CHOE et al. within the combined modified teaching of the Computerized machine learning of interesting video sections mentioned by Suri et al. and the Learning to Detect Multiple Photographic Defects mentioned by Ning Yu et al. provides a system and method for implementing machine learning algorithm for image processing.
It is noted that any citations to specific, pages, columns, lines, or figures in the prior art references and any interpretation of the reference should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. See MPEP 2123.
Allowable Subject Matter
4. Claims 4,6,11,12,16,17,18 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
5. The following is an examiner’s statement of reasons for objecting the claims as allowable subject matter:
As to claims 4 and 11, prior art of record does not teach or suggest the limitation mentioned within claims 4 and 11 : “…a bottom layer of the first neural network is trained to classify the image based on whether the image includes the exposure defect; and a bottom layer of the second neural network is trained to classify the image based on the level of exposure.”
As to claim 6, prior art of record does not teach or suggest the limitation mentioned within claim 6 : “…the first neural network is trained using a weak supervised learning algorithm, and the training of the binary model determines an inference score using the dataset of images, the dataset of images being a noisy dataset.”
As to claim 12, prior art of record does not teach or suggest the limitation mentioned within claim 12 : “…a learning rate used when training the first neural network and the second neural network is lower at the top layer relative to the separate bottom layers.”
As to claim 16, prior art of record does not teach or suggest the limitation mentioned within claim 16 : “…the first neural network binary model and the second neural network classification model share a top layer; and the first neural network binary model and the second neural network classification model each have a separate bottom layer.”
As to claims 17 and 18, the following claim 17 and 18 depend on objected allowable claim 16, therefore claims 17 and 18 are objected as allowable claim over prior art on record.
As to claim 19, prior art of record does not teach or suggest the limitation mentioned within claim 19 : “…the first neural network binary model is trained using a weak supervised learning algorithm, and the training of the first neural network binary model determines an inference score using the dataset of images.”
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Examiner’s Notes
5. The Examiner acknowledges the following prior arts below as pertinent to the current applications claim limitations and inventive concept, although the following prior arts shown below were not relied upon to address the limitations within the claim , they are analogous art mentioning the inventive concept key points on ( Image processing, exposure defect classification, neural network models etc. ).
1) Jing Wang et al., “Exposure correction using deep learning,”May 13th 2019,Journal of Electronic Imaging,Vol. 28(3), 033003 (2019),Pages 033003-01-11.
2) K. Ram Prabhakar et al., "DeepFuse: A Deep Unsupervised Approach for Exposure Fusion with Extreme Exposure Image Pairs," October ,2017, Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, Pages 4714-4721.
3) Chuanfei Hu et al.,"An Efficient Convolutional Neural Network Model Based on Object-Level Attention Mechanism for Casting Defect Detection on Radiography Images," 18th August 2020, IEEE TRANSACTIONS ON INDUSTRIAL ELECTRONICS, VOL. 67, NO. 12, DECEMBER 2020,Pages 10922-10928.
4) Xiaopeng Wang et al.,"Binary classification of welding defect based on deep learning," 12th March 2022, SCIENCE AND TECHNOLOGY OF WELDING AND JOINING 2022, VOL. 27, NO. 6,Pages 407-416.
5) Cristiano R. Steffens et al.,"CNN Based Image Restoration,"11th January 2020, Journal of Intelligent & Robotic Systems (2020),Pages 609-624.
6) Weihuang Liu et al.,"Convolutional Two-Stream Network UsingMulti-Facial Feature Fusion for Driver Fatigue Detection," 14th May 2019, Future Internet 2019, 11, 115,Page 1- 11.
7) Husein Perez et al.,"Deep Learning for Detecting Building Defects Using Convolutional Neural Networks,"15th August 2019, Sensors 2019, 19, 3556,Pages 1-18.
8) Cristiano Rafael Steffens et al.,"Deep Learning based Exposure Correction for Image Exposure Correction with Application in Computer Vision for Robotics," 27th December 2018 Latin American Robotic Symposium, 2018 Brazilian Symposium on Robotics (SBR) and 2018 Workshop on Robotics in Education (WRE) Pages 194-199.
9) Xinghui Dong et al., "Defect Detection and Classification by Training a Generic Convolutional Neural Network Encoder," 29th October 2020, IEEE TRANSACTIONS ON SIGNAL PROCESSING, VOL. 68, 2020,Pages 6055-6067.
10) Jianrui Cai et al., "Learning a Deep Single Image Contrast Enhancer from Multi-Exposure Images," 6th February 2018, IEEE TRANSACTIONS ON IMAGE PROCESSING, VOL. 27, NO. 4, APRIL 2018.Page 2049-2060.
11) Qing Zhang et al.,"High-Quality Exposure Correction of Underexposed Photos," 15th October 2018, MM '18: Proceedings of the 26th ACM international conference on Multimedia,Pages 582-588.
12) Mahmoud Afifi et al.,"Learning to Correct Overexposed and Underexposed Photos," 25th March 2020, arXiv:2003.11596v1, Pages 1-17.
13) Steinberg et al. (USPUB 20080317357 )
14) Steinberg et al. (USPUB 20080316328 )
15) Wang et al. (USPUB 20190333198 )
16) Li et al. (USPUB 20200265153 )
17) CHEN Q et al. ( CN 112233066 )
18) BHARTI et al. (USPUB 20210117729 )
19) Kumar et al. (USPUB 20220253990)
20) RITTSCHER et al. (USPUB 20220207728 )
21) Kumar et al. (USPUB 20170111137 )
22) Xu et al.( USPUB 20230003838 )
23) Cherubini et al. ( USPUB 20230110263 )
24) Blais-Morin et al. (USPUB 20230112788 )
Conclusion
6. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of Reference Cited for a listing of analogous art.
7. Any inquiry concerning this communication or earlier communications from the examiner should be directed to OMAR S ISMAIL whose telephone number is (571)272-9799 and Fax # is (571)273-9799. The examiner can normally be reached on M-F 9:00am-6:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at
http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David C. Payne can be reached on (571) 272-3024. The fax phone number for the organization where this application or proceeding is assigned is (571)273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free)? If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/OMAR S ISMAIL/
Primary Examiner, Art Unit 2635