Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see Applicants Remarks pages 6-9, filed 08/19/26, with respect to the rejection(s) of claim(s) 1-20 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Chung et al US 20210027065.
Regarding claim 1, 11 and 16, Applicant states that Tajbakhsh, Li and Marwah fails to teach a second pretrained machine learning model that generates a single classification for a target feature represented by a plurality of embedding vectors, wherein the single classification is generated from a joint analysis of the plurality of embedding vectors (Applicants Remarks pages 6-7).
First, Applicant states that Tajbakhsh fails to teach the claimed architecture in which a first pretrained machine learning model generates a plurality of embedding vectors that are then used by a second pretrained machine learning model to classify a target feature and that Marwah merely discloses jointly learning embedded vectors during model training rather than generating a classification from the joint analysis of the plurality of embedding vectors (Applicants Remarks page 7). Examiner disagrees with Applicant. Tajbakhsh teaches each set of generated patches may be fed to the corresponding deep CNNs, with each patch being assigned a probabilistic score. The maximum probabilistic score for each set of patches may then be computed by the processor. A global probabilistic score for the polyp candidate may then be computed by the processor by averaging the maximum probabilistic scores for each set of patches (column 6, lines 30-36). Because each set of patches are fed into corresponding CNNs, this would read on a first pretrained machine learning model would read on one of the CNNs that correspond to a set of patches, since a second set of patches would be fed to a second set of CNNS, etc. The image patches of Tajbakhsh are read as embedding vector. A patch-based CNN (Tajbakhsh CNNS) divides the image into smaller, non-overlapping (or overlapping) patches. Each patch is processed independently, often with its own convolutional filters. The output of a patch can be a local embedding vector — for example, the result of a small convolutional layer applied to that patch (https://www.codegenes.net/blog/patch-based-cnn-pytorch/)
Marwah teaches of machine learning model may learn first embedded vectors and second embedded vectors. The processor may execute the instructions to learn the first embedded vectors and the second embedded vectors jointly. (paragraph 0078). Marwah is not relied on to teach of generating classification from a joint analysis of a plurality of embedding vectors. The teaching of joint analysis of first and second embedding vectors of Marwah can be applied in combination, so that the machine learning model of Chung et al can analyze a plurality of embedding vectors from a video and the image patches (which read on embedding vector) CNN learning model of Tajbakhsh (first pretrained learning model
Second, Applicant states that the reference of Li fails to generate a classification of a target feature using a plurality of embedding vectors generated from a plurality of frames of a medical video (Applicants Remarks pages 7-8). Examiner agrees with Applicant.
Chung et al teaches generating, by the second pretrained machine learning model, a classification of the target feature using the plurality of embedding vectors (model training module 204 can train a machine learning model based on the labeled set of training data. Each video in the set of training videos (and/or the set of qualified training videos) can be associated with a set of video features. the machine learning model comprises a deep neural network cascaded with a sparse neural network. sparse neural network model can be trained to receive other video features (e.g., video metadata, video creator features), as well as the embedding (e.g., vector representation) of the video generated by the deep neural network model and generate a final output pertaining to video quality for a video. the final output may be an overall video quality score indicative of a quality of the video (classification) (paragraph 0040) Note: The embedding is per video in the set of training videos, thereby disclosing a plurality of embedding vectors to train machine learning model and generate a classification of a video feature,
wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, (model training module 204 can train a machine learning model based on the labeled set of training data. As discussed above, each video in the set of training videos (and/or the set of qualified training videos) can be associated with a set of video features. The model can be trained, based on the labels and the video features, to identify which video features are most likely to result in a “high quality video” as defined by the identified evaluation objective (jointly analyzing plurality of embedding vectors to each frame since a video consists of frames, see paragraph 0036). sparse neural network model (which is apart of model training module 204) can be trained to receive other video features (e.g., video metadata, video creator features), as well as the embedding (e.g., vector representation) of the video generated by the deep neural network model and generate a final output pertaining to video quality for a video. the final output may be an overall video quality score indicative of a quality of the video (paragraph 0040) Note: the video quality is a classification of the target feature using video features and embedding of the set of training videos
the single classification being generated from the joint analysis of the plurality of embedding vectors (model training module 204 can be configured to train a machine learning model using training data. the training data can include a set of training videos. Each video of the set of training videos can be associated with a set of video features. the set of video features and image information associated with the video (e.g., image content of the video, objects identified in the video frames (paragraph 0036). the machine learning model can be trained to receive image and sound data associated with a video (video contains frames), and generate an embedding (i.e., an n-dimensional vector representation) of the video. Note: embedding is generated for each video in the set of training videos, thereby reading on a plurality of embedding vectors. The model can be trained, based on the labels and the video features, to identify which video features are most likely to result in a “high quality video” as defined by the identified evaluation objective (jointly analyzing plurality of embedding vectors to each frame since a video consists of frames) and generate a final output pertaining to video quality for a video. the final output may be an assignment of the video to a particular video quality category of a pre-defined set of video quality categories (e.g., poor, fair, good, excellent) indicative of a likelihood of the video to achieve an associated video quality metric requirement (classification) (paragraph 0040)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-8 and 10-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tajbakhsh et al US 10055843 in view of Chung et al US 20210027065 further in view of Marwah et al US 20200394463.
Regarding claim 1, Tajbakhsh et al teaches a method of classifying a target feature in a medical video by one or more computer systems (the polyp detection system 100 to detect colonic polyps during a colonoscopy procedure (column 4, lines 1-8), wherein the one or more computer systems comprises a first pretrained machine learning model and a second pretrained machine learning model (a set of CNNs (first pretrained machine learning model and second pretrained machine learning model) are applied to the corresponding image patches (embedding vector), and probabilities indicative of a maximum response for each CNN are computed, as indicated by process block 410. (column 8, lines 16-35) CNNs, a polyp detection method. (column 9, lines 56-60), the method comprising:
receiving a plurality of frames of the medical video, wherein the plurality of frames comprises the target feature (A set of 40 short colonoscopy videos were collected, of which half were positive and half were negative shots. A positive shot was defined as a sequence of frames that showed a unique polyp (target feature) from different view angles (column 9, lines 40-43);
generating, by the first pretrained machine learning model, an embedding vector for each frame of the plurality of frames, each embedding vector having a predetermined number of values (a set of CNNs (first pretrained machine learning model) are applied to the corresponding image patches (embedding vector), and probabilities indicative of a maximum response for each CNN are computed, as indicated by process block 410. That is, each patch is assigned a probabilistic score (predetermined number of values). The maximum probabilistic score for each set of patches may then be computed, and a global probabilistic score may then be computed by averaging the maximum probabilistic scores for each set of patches. In this manner, a confidence value for each polyp candidate can be generated (column 8, lines 16-35) Note: the patches are collected from frames of polyps identified (column 8, lines 64-67); and,
Although Tajbakhsh et al teaches train the CNNs, a polyp detection method using a CVCColonDB database was applied, similar to previous work done by the inventors. All the generated polyp candidates were grouped into true and false detections according to the available ground truth for the training videos (column 9, lines 56-65)
Tajbakhsh et al fails to teach generating, by the second pretrained machine learning model, a classification of the target feature using the plurality of embedding vectors, wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors
Chung et al teaches generating, by the second pretrained machine learning model, a classification of the target feature using the plurality of embedding vectors (model training module 204 can train a machine learning model based on the labeled set of training data. Each video in the set of training videos (and/or the set of qualified training videos) can be associated with a set of video features. the machine learning model comprises a deep neural network cascaded with a sparse neural network. sparse neural network model can be trained to receive other video features (e.g., video metadata, video creator features), as well as the embedding (e.g., vector representation) of the video generated by the deep neural network model and generate a final output pertaining to video quality for a video. the final output may be an overall video quality score indicative of a quality of the video (classification) (paragraph 0040) Note: The embedding is per video in the set of training videos, thereby disclosing a plurality of embedding vectors to train machine learning model and generate a classification of a video feature,
wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, (model training module 204 can train a machine learning model based on the labeled set of training data. As discussed above, each video in the set of training videos (and/or the set of qualified training videos) can be associated with a set of video features. The model can be trained, based on the labels and the video features, to identify which video features are most likely to result in a “high quality video” as defined by the identified evaluation objective (jointly analyzing plurality of embedding vectors to each frame since a video consists of frames, see paragraph 0036). sparse neural network model (which is apart of model training module 204) can be trained to receive other video features (e.g., video metadata, video creator features), as well as the embedding (e.g., vector representation) of the video generated by the deep neural network model and generate a final output pertaining to video quality for a video. the final output may be an overall video quality score indicative of a quality of the video (paragraph 0040) Note: the video quality is a classification of the target feature using video features and embedding of the set of training videos
the single classification being generated from the joint analysis of the plurality of embedding vectors (model training module 204 can be configured to train a machine learning model using training data. the training data can include a set of training videos. Each video of the set of training videos can be associated with a set of video features. the set of video features and image information associated with the video (e.g., image content of the video, objects identified in the video frames (paragraph 0036). the machine learning model can be trained to receive image and sound data associated with a video (video contains frames), and generate an embedding (i.e., an n-dimensional vector representation) of the video. Note: embedding is generated for each video in the set of training videos, thereby reading on a plurality of embedding vectors. The model can be trained, based on the labels and the video features, to identify which video features are most likely to result in a “high quality video” as defined by the identified evaluation objective (jointly analyzing plurality of embedding vectors to each frame since a video consists of frames) and generate a final output pertaining to video quality for a video. The final output may be an assignment of the video to a particular video quality category of a pre-defined set of video quality categories (e.g., poor, fair, good, excellent) indicative of a likelihood of the video to achieve an associated video quality metric requirement (classification) (paragraph 0040).
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al to include: generating, by the second pretrained machine learning model, a classification of the target feature using the plurality of embedding vectors, wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors.
The reason for doing so would be to accurately identify objects in an image or video.
Tajbakhsh et al in view of Chung et al fails to teach wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the medical video
Marwah et al teaches wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors (processor may train a machine learning model machine learning model may learn first embedded vectors and second embedded vectors. The processor may execute the instructions to learn the first embedded vectors and the second embedded vectors jointly. (paragraph 0078). It is well know that machine learning models can be trained to perform any process as taught. Therefore, the teaching of Marwah et al can be applied to the plurality of frames taught in Tajbakhsh et al)
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al to include: wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors.
The reason for doing so would be to save processing time and perform analysis faster.
Regarding claim 3, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification comprises a score, wherein the score is in a range of 0 to 1 (Tajbakhsh et al: each of the trained CNNs significantly improved performance compared to the previous method (p<0:0001, JAFROC test), the best result was obtained using score fusion framework, with the number of false positives being reduced by a factor of 50 at 50% sensitivity and by a factor of 22 at 60% sensitivity. The score fusion approach generated 0.002 false positives per frame at 50% sensitivity, a significant performance improvement compared to a previous technique, which produced 0.15 false positives per frame at the same sensitivity. (column 10, lines 53-65) Note: the score is based on positive or false positive classification.
Regarding claim 4, Tajbakshh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification is selected from positive, negative, and uncertain (Tajbakhsh et al: For evaluation, a detection was considered as a true (false) positive if it fell inside (outside) the white region of the ground truth image. (column 9, lines 52-55).
Regarding claim 5, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification comprises a textual representation (Tajbakhsh et al: A report is then generated, at process block 414. As described, the report may provide audio and/or visual information. For instance, raw or processed optical images may be displayed, along with indicators and locations for identified objects, such as polyps, vessels, lumens, specular reflections, and so forth. The report may also indicate the probabilities or confidence scores for identified objects, including colonic polyps (column 8, lines 38-45) Note: the polyps are identified and classified using CNNs and the result of the classification is generated in the report with text describing the identified object ).
Regarding claim 6, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the first pretrained machine learning model and the second pretrained machine learning model are jointly trained (Tajbakhsh et al: each of the trained CNNs significantly improved performance (column 10, lines 53-65).
Regarding claim 8, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the medical video is collected during a colonoscopy procedure using an endoscope and wherein the target feature is a polyp (Tajbakhsh et al : the polyp detection system 100 may generally include a colonoscopy device 102. the colonoscopy device 102 and controller 104 may be utilized to detect colonic polyps during a colonoscopy procedure (column 4, lines 1-8). the colonoscopy device 102 may include an endoscope (not shown in FIG. 1) configured to acquire optical image data, either continuously or intermittently, from a patient's colon, and relay the optical image data to the controller 104 for processing and analysis (column 4, lines 9-11) .
Regarding claim 11, Tajbakhsh et al teaches a system for classifying a target feature in a medical video (the polyp detection system 100 to detect colonic polyps during a colonoscopy procedure (column 4, lines 1-8) comprising:
an input interface configured to receive a medical video (the polyp detection system 100 may generally include a colonoscopy device 102. the colonoscopy device 102 and controller 104 may be utilized to detect colonic polyps during a colonoscopy procedure (column 4, lines 1-8). data received by the input elements includes optical image data, or video data obtained during a colonoscopy. (column 5, lines 31-34);
a memory configured to store a plurality of processor-executable instructions (a data storage or memory (column 5, lines 47-50), the memory including:
an embedder based on a first pretrained machine learning model (a set of CNNs (first pretrained machine learning model/embedder) are applied to the corresponding image patches , and probabilities indicative of a maximum response for each CNN are computed, as indicated by process block 410. (column 8, lines 16-35); and,
a processor configured to execute the plurality of processor-executable instruction to perform operations including:
receiving a plurality of frames of the medical video, wherein the plurality of frames comprises the target feature (A set of 40 short colonoscopy videos were collected, of which half were positive and half were negative shots. A positive shot was defined as a sequence of frames that showed a unique polyp (target feature) from different view angles (column 9, lines 40-43);
generating, with the embedder, an embedding vector for each frame of the plurality of frames, each embedding vector having a predetermined number of values (a set of CNNs (first pretrained machine learning model) are applied to the corresponding image patches (embedding vector), and probabilities indicative of a maximum response for each CNN are computed, as indicated by process block 410. That is, each patch is assigned a probabilistic score (predetermined number of values). The maximum probabilistic score for each set of patches may then be computed, and a global probabilistic score may then be computed by averaging the maximum probabilistic scores for each set of patches. In this manner, a confidence value for each polyp candidate can be generated (column 8, lines 16-35); and,
Although Tajbakhsh et al teaches train the CNNs, a polyp detection method using a CVCColonDB database was applied, similar to previous work done by the inventors. All the generated polyp candidates were grouped into true and false detections according to the available ground truth for the training videos (column 9, lines 56-65)
Tajbakhsh et al fails to teach a classifier based on a second pretrained machine learning model; and,
generating, with the classifier, a classification of the target feature using the plurality of embedding vectors, wherein the classifier jointly analyzes the plurality of embedding vectors jointly corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors.
Chung et al teaches a classifier based on a second pretrained machine learning model (The sparse neural network model can be trained to receive other video features (e.g., video metadata, video creator features), as well as the embedding (e.g., vector representation) of the video generated by the deep neural network model, and generate a final output pertaining to video quality for a video. For example, the final output may be an overall video quality score indicative of a quality of the video. In another example, the final output may be an assignment of the video to a particular video quality category of a pre-defined set of video quality categories (e.g., poor, fair, good, excellent) indicative of a likelihood of the video to achieve an associated video quality metric requirement (classification) (paragraph 0040); and,
generating, with the classifier, a classification of the target feature using the plurality of embedding vectors, wherein the classifier jointly analyzes the plurality of embedding vectors jointly corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors (model training module 204 can be configured to train a machine learning model using training data. In various embodiments, the training data can include a set of training videos. Each video of the set of training videos can be associated with a set of video features. In certain embodiments, the set of video features and image information associated with the video (e.g., image content of the video, objects identified in the video frames (paragraph 0036). the machine learning model can be trained to receive image and sound data associated with a video (video contains frames), and generate an embedding (i.e., an n-dimensional vector representation) of the video. The sparse neural network model can be trained to receive other video features (e.g., video metadata, video creator features), as well as the embedding (e.g., vector representation) of the video generated by the deep neural network model, and generate a final output pertaining to video quality for a video. For example, the final output may be an overall video quality score indicative of a quality of the video. In another example, the final output may be an assignment of the video to a particular video quality category of a pre-defined set of video quality categories (e.g., poor, fair, good, excellent) indicative of a likelihood of the video to achieve an associated video quality metric requirement (classification) (paragraph 0040).
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al to include: a classifier based on a second pretrained machine learning model; and,
generating, with the classifier, a classification of the target feature using the plurality of embedding vectors, wherein the classifier jointly analyzes the plurality of embedding vectors jointly corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors.
The reason for doing so would be to accurately identify objects in an image or video.
Tajbakhsh et al in view of Chung et al fails to teach wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the medical video
Marwah et al teaches wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors (processor may fetch, decode, and execute the instructions 604 to train a machine learning model. machine learning model may learn first embedded vectors and second embedded vectors. The processor may execute the instructions to learn the first embedded vectors and the second embedded vectors jointly. (paragraph 0078). It is well know that machine learning models can be trained to perform any process as taught. Therefore, the teaching of Marwah et al can be applied to the plurality of frames taught in Tajbakhsh et al)
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al to include: wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors.
The reason for doing so would be to save processing time and perform analysis faster.
Regarding claim 13, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification comprises a score, wherein the score is in a range of 0 to 1 (Tajbakhsh et al: each of the trained CNNs significantly improved performance compared to the previous method (p<0:0001, JAFROC test), the best result was obtained using score fusion framework, with the number of false positives being reduced by a factor of 50 at 50% sensitivity and by a factor of 22 at 60% sensitivity. The score fusion approach generated 0.002 false positives per frame at 50% sensitivity, a significant performance improvement compared to a previous technique, which produced 0.15 false positives per frame at the same sensitivity. (column 10, lines 53-65) Note: the score is based on positive or false positive classification.
Regarding claim 14, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification comprises one of: positive, negative, or uncertain (Tajbakhsh et al: For evaluation, a detection was considered as a true (false) positive if it fell inside (outside) the white region of the ground truth image. (column 9, lines 52-55).
Regarding 15, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification comprises a textual representation (Tajbakhsh et al: A report is then generated, at process block 414. As described, the report may provide audio and/or visual information. For instance, raw or processed optical images may be displayed, along with indicators and locations for identified objects, such as polyps, vessels, lumens, specular reflections, and so forth. The report may also indicate the probabilities or confidence scores for identified objects, including colonic polyps (column 8, lines 38-45) Note: the polyps are identified and classified using CNNs and the result of the classification is generated in the report with text describing the identified object ).
Regarding claim 16, Tajbakhsh et al teaches a non-transitory processor-readable storage medium storing a plurality of processor-executable instructions (the processor may read and execute software instructions from a non-transitory computer-readable medium, (column 5, lines 47-52) for classifying a target feature in a medical video ( (the polyp detection system 100 to detect colonic polyps during a colonoscopy procedure (column 4, lines 1-8), the plurality of processor executable instructions being executed by a processor to perform operations comprising:
receiving a plurality of frames of the medical video, wherein the plurality of frames comprises the target feature (A set of 40 short colonoscopy videos were collected, of which half were positive and half were negative shots. A positive shot was defined as a sequence of frames that showed a unique polyp (target feature) from different view angles (column 9, lines 40-43);
generating, by a first pretrained machine learning model, an embedding vector for each frame of the plurality of frames, each embedding vector having a predetermined number of values (a set of CNNs (first pretrained machine learning model) are applied to the corresponding image patches (embedding vector), and probabilities indicative of a maximum response for each CNN are computed, as indicated by process block 410. That is, each patch is assigned a probabilistic score (predetermined number of values). The maximum probabilistic score for each set of patches may then be computed, and a global probabilistic score may then be computed by averaging the maximum probabilistic scores for each set of patches. In this manner, a confidence value for each polyp candidate can be generated (column 8, lines 16-35); and,
Although Tajbakhsh et al teaches train the CNNs, a polyp detection method using a CVCColonDB database was applied, similar to previous work done by the inventors. All the generated polyp candidates were grouped into true and false detections according to the available ground truth for the training videos (column 9, lines 56-65)
Tajbakhsh et al fails to teach generating, by the second pretrained machine learning model, a classification of the target feature using the plurality of embedding vectors, wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors
Chung et al teaches generating, by the second pretrained machine learning model, a classification of the target feature using the plurality of embedding vectors, wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors (model training module 204 can be configured to train a machine learning model using training data. In various embodiments, the training data can include a set of training videos. Each video of the set of training videos can be associated with a set of video features. In certain embodiments, the set of video features and image information associated with the video (e.g., image content of the video, objects identified in the video frames (paragraph 0036). the machine learning model can be trained to receive image and sound data associated with a video (video contains frames), and generate an embedding (i.e., an n-dimensional vector representation) of the video. The sparse neural network model can be trained to receive other video features (e.g., video metadata, video creator features), as well as the embedding (e.g., vector representation) of the video generated by the deep neural network model, and generate a final output pertaining to video quality for a video. For example, the final output may be an overall video quality score indicative of a quality of the video. In another example, the final output may be an assignment of the video to a particular video quality category of a pre-defined set of video quality categories (e.g., poor, fair, good, excellent) indicative of a likelihood of the video to achieve an associated video quality metric requirement (classification) (paragraph 0040).
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al to include: generating, by the second pretrained machine learning model, a classification of the target feature using the plurality of embedding vectors, wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the video to generate a single classification for the target feature represented by the plurality of embedding vectors, the single classification being generated from the joint analysis of the plurality of embedding vectors.
The reason for doing so would be to accurately identify objects in an image or video.
Tajbakhsh et al in view of Chung et al fails to teach wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors corresponding to each frame of the plurality of frames in the medical video
Marwah et al teaches wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors (processor may fetch, decode, and execute the instructions 604 to train a machine learning model. machine learning model may learn first embedded vectors and second embedded vectors. The processor may execute the instructions to learn the first embedded vectors and the second embedded vectors jointly. (paragraph 0078). It is well know that machine learning models can be trained to perform any process as taught. Therefore, the teaching of Marwah et al can be applied to the plurality of frames taught in Tajbakhsh et al)
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al to include: wherein the second pretrained machine learning model jointly analyzes the plurality of embedding vectors.
The reason for doing so would be to save processing time and perform analysis faster.
Regarding claim 18, Tajbakhsh et al in view of Li et al teaches wherein the classification comprises a score, wherein the score is in a range of 0 to 1 (Tajbakhsh et al: each of the trained CNNs significantly improved performance compared to the previous method (p<0:0001, JAFROC test), the best result was obtained using score fusion framework, with the number of false positives being reduced by a factor of 50 at 50% sensitivity and by a factor of 22 at 60% sensitivity. The score fusion approach generated 0.002 false positives per frame at 50% sensitivity, a significant performance improvement compared to a previous technique, which produced 0.15 false positives per frame at the same sensitivity. (column 10, lines 53-65) Note: the score is based on positive or false positive classification.
Regarding claim 19, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification comprises one of: positive, negative, or uncertain (Tajbakhsh et al: For evaluation, a detection was considered as a true (false) positive if it fell inside (outside) the white region of the ground truth image. (column 9, lines 52-55).
Regarding claim 20, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches wherein the classification comprises a textual representation (Tajbakhsh et al: A report is then generated, at process block 414. As described, the report may provide audio and/or visual information. For instance, raw or processed optical images may be displayed, along with indicators and locations for identified objects, such as polyps, vessels, lumens, specular reflections, and so forth. The report may also indicate the probabilities or confidence scores for identified objects, including colonic polyps (column 8, lines 38-45) Note: the polyps are identified and classified using CNNs and the result of the classification is generated in the report with text describing the identified object ).
Claim(s) 2, 7, 10, 12 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tajbakhsh et al US 10055843 in view of Chung et al US 20210027065 further in view of Marwah et al US 20200394463 further in view of Li et al US 20210201701.
Regarding claim 2, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches all of the limitations of claim 1
Tajbakhsh et al teaches wherein the first pretrained learning model comprises a convolutional neural network (CNNs, a polyp detection method. (column 9, lines 56-60), and
Tajbakhsh et al in view of Chung et al further in view of Marwah et al fails to teach wherein the second pretrained machine learning model comprises a transformer.
Li et al teaches wherein the second pretrained machine learning model comprises a transformer (second server 120 may also extract contextual semantic feature(s) from the vector or the sequence of vectors using, for example, a recurrent neural network (RNN), a convolutional neural network (CNN), a transformer, or the like (paragraph 0061).
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al further in view of Marwah et al to include: wherein the second pretrained machine learning model comprises a transformer.
The reason for doing so would be to accurately identify vectors in an image or video.
Regarding claim 7, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches all of the limitations of claim 1
Tajbakhsh et al in view of Chung et al further in view of Marwah et al fails to teach wherein the first pretrained machine learning model and the second pretrained machine learning model are trained separately
Li et al teaches wherein the first pretrained machine learning model and the second pretrained machine learning model are trained separately (The first server(s) 110 may be configured to obtain and/or generate training material used in medical diagnosis training. The training material may relate to one or more medical images of one or more subjects (paragraph 0053). The difficulty level of a medical image may be obtained from a first server that collects the medical image and/or determined by the second server (paragraph 0061) Note: the first server and second server operate as machine learning models. The first server uses training material and the second server comprises RNN, CNN, etc (paragraph 0061)
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al further in view of Marwah et al to include: wherein the first pretrained machine learning model and the second pretrained machine learning model are trained separately
The reason for doing so would be to process the video frames of a video faster.
Regarding claim 10, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches all of the limitations of claim 1
Tajbakhsh et al in view of Chung et al further in view of Marwah et al fails to teach wherein the second pretrained machine learning model analyzes the plurality of embedding vectors without classifying each embedding vector individually
Li et al teaches wherein the second pretrained machine learning model analyzes the plurality of embedding vectors without classifying each embedding vector individually (the second server 120 may transform the reference diagnostic result of the medical image into a vector or a sequence of vectors using a word embedding technique. The second server 120 may also extract contextual semantic feature(s) from the vector or the sequence of vectors using, for example, a recurrent neural network (RNN), a convolutional neural network (CNN), along short term memory (LSTM) model, a transformer, or the like (paragraph 0061).
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al further in view of Marwah et al to include: wherein the second pretrained machine learning model analyzes the plurality of embedding vectors without classifying each embedding vector individually
The reason for doing so would be to process the video frames of a video faster.
Regarding claim 12, Tajbakhsh et al teaches wherein the first pretrained learning model comprises a convolutional neural network (CNNs, a polyp detection method. (column 9, lines 56-60), and
Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches all of the limitations of claim 11 fails to teach wherein the second pretrained machine learning model comprises a transformer.
Li et al teaches wherein the second pretrained machine learning model comprises a transformer (second server 120 may also extract contextual semantic feature(s) from the vector or the sequence of vectors using, for example, a recurrent neural network (RNN), a convolutional neural network (CNN), a transformer, or the like (paragraph 0061).
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al further in view of Marwah et al to include: wherein the second pretrained machine learning model comprises a transformer.
The reason for doing so would be to accurately identify vectors in an image or video.
Regarding claim 17, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches all of the limitations of claim 16
Tajbakhsh et al teaches wherein the first pretrained learning model comprises a convolutional neural network (CNNs, a polyp detection method. (column 9, lines 56-60), and
Tajbakhsh et al in view of Chung et al further in view of Marwah et al fails to teach wherein the second pretrained machine learning model comprises a transformer.
Li et al teaches wherein the second pretrained machine learning model comprises a transformer (second server 120 may also extract contextual semantic feature(s) from the vector or the sequence of vectors using, for example, a recurrent neural network (RNN), a convolutional neural network (CNN), a transformer, or the like (paragraph 0061).
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al further in view of Marwah et al to include: wherein the second pretrained machine learning model comprises a transformer.
The reason for doing so would be to accurately identify vector in an image or video.
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tajbakhsh et al US 10055843 in view of Chung et al US 20210027065 further in view of Marwah et al US 20200394463 further in view of Veidman et al US 20200372635.
Regarding claim 9, Tajbakhsh et al in view of Chung et al further in view of Marwah et al teaches all of the limitation of claims 1 and 8
Tajbakhsh et al in view of Chung et al further in view of Marwah et al fails to teach wherein the classification comprises one of: adenomatous and non-adenomatous.
Veidman et al teaches wherein the classification comprises one of: adenomatous and non-adenomatous (the received indication of polyp, exemplary potential slide-level type type(s) include: hyperplastic, low grade adenoma, high grade adenoma, and malignant. Optionally, the slide-level tissue type(s) include slide-level sub-tissue type(s), for example, for slide-level tissue type of Adenomatous Lesion, potential sub-tissue type(s) include tubular, tubulovillous, and villous (paragraph 0201)
Therefore, it would have been obvious to a person of ordinary skill in the art to modify Tajbakhsh et al in view of Chung et al further in view of Marwah et al to include: wherein the classification comprises one of: adenomatous and non-adenomatous.
The reason for doing so would be to accurately identify the type of polyp.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL L BURLESON whose telephone number is (571)272-7460.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Akwasi Sarpong can be reached on 571 270-3438 The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Michael Burleson
Patent Examiner
Art Unit 2683
Michael Burleson
September 14, 2026
/MICHAEL BURLESON/
/AKWASI M SARPONG/SPE, Art Unit 2681 9/21/2026