DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application is being examined under the pre-AIA first to invent provisions.
Claims 1-20 are pending and examined below. The application is a 35 USC 371 national stage entry of PCT/US2023/024902 filed June 9, 2023, which claims benefit of provisional applications 63/458,054 filed April 7, 2023, and 63/350,720 filed June 9, 2022. Provisional 63/350,720 provides support for the claims and cites Appendix B and Appendix C, two of the Zamzmi papers cited on the IDS that provide support for the June 9, 2022, effective filling date.
Information Disclosure Statement
The IDS(s) has/have been considered and placed in the application file.
Claim Objections
Claim 5 is objected to because of the following informality. Claim 5 recites “a spatial
transformation network to focus on specific regions the selected images.” A word appears to
have been omitted between “regions” and “the.” Appropriate correction is required. Claims 6-8 are objected to by dependence on claim 5.
Claims 18-20 are objected to as being in improper dependent form. Claim 18 recites “The method according to claim 16,” and claims 19 and 20 depend in turn from claims 18 and 19, but claim 16 is directed to “A non-transitory computer readable medium” and not to a
method. Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 11 and 16 are rejected under 35 U.S.C. 112(6) as being indefinite. Each recites
finding a diameter “at different spatial (different sites) and temporal points (over time).” It
cannot be determined whether the parentheticals limit the claim, define the terms they follow, are exemplary. MPEP 2173.05(d). Claims 2-10, 12-15 and 17-20 depend from claims 1, 11 and
16, and are also rejected by dependence.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 11-14 and 16-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim(s) recite(s) a process (claim 11) and an article of manufacture (claim 16), which is one of the four statutory categories of invention.
Step 2A, Prong One: The claim recites receiving images from the echocardiography study and performing a quality assessment of the images to generate selected images; localizing and segmenting the selected images to obtain IVC regions in the selected images; performing a distance calculation to find a diameter in the IVC regions at different spatial and temporal points; classifying known views and unknown views in the selected images; and grouping the unknown views into one or more clusters used to update classification capability in the AI system.
These limitations fall within the mental processes (concepts performed in the human mind, including observation, evaluation, judgment, and opinion) grouping of abstract ideas, because a sonographer judges echo image quality by eye, identifies the inferior vena cava by eye, and sorts views into known and unknown by observation and judgment, all of which can be performed in the human mind.
Step 2A, Prong Two: The claim recites the following additional element(s): an Artificial
Intelligence (AI) system (claim 11), and a non-transitory computer readable medium comprising
instructions executed by a computer ( claim 16). The additional elements are the recitation of an Artificial Intelligence system and receiving images from the echocardiography study. The AI
system is recited at the highest level of generality, with no architecture, no layers, no encoder, no pathway, and no clustering algorithm. The AI system is a generic computer invoked as a tool to perform the exception. Receiving the images is insignificant extra-solution data gathering. The additional element(s) therefore do not integrate the judicial exception into a practical application.
Step 2B: The claim does not include additional elements that are sufficient to amount to
significantly more than the judicial exception.
The additional element(s) of an Artificial Intelligence (Al) system is/are described in the
specification at ¶ 48 as “Existing systems for automated echocardiography analysis are designed under the assumption that the examples in the testing or deployment stage must belong to the same limited number of classes that have appeared in the training stage,” which is a statement demonstrating the well-understood, routine, and conventional nature of the element. MPEP 2106.05(d)(I).
The additional element(s) of a non-transitory computer readable medium comprising instructions executed by a computer amount(s) to storing and retrieving information in memory, which the courts have recognized as well-understood, routine, and conventional. MPEP 2106.05(d)(II).
The claim does no more than apply generic machine learning to a new data environment. Claims that “do no more than claim the application of generic machine learning to new data environments, without disclosing improvements to the machine learning models to be applied” are not eligible. Recentive Analytics, Inc. v. Fox Corp., 134 F.4th 1205 (Fed. Cir. 2025).
By contrast, claims directed to improving the functioning of the model itself may be eligible. See MPEP 2106.04(d), subsection III, and MPEP 2106.05(a), citing Ex parte Desjardins, Appeal No. 2024-000567 (Appeals Review Panel Sept. 26, 2025). No such improvement to the model is recited here; the specification identifies the improvement as a 55,620 parameter retrieval network, segmentation at 85 frames per second, and a trilateral pathway architecture, while claims 11 and 16 recite only the outcome, selected images and segmented regions and a diameter and clusters, without the mechanism that produces it. Claims 1-10 integrate the exception into a practical application and are not rejected under 35 U.S.C. 101 because the 112(f) limitations correspond to a shared-encoder two-head retrieval network (Spec ¶¶ 22-23), a segmentation network with a spatial transformer and thee pathways (Spec ¶¶ 39-45), and a contour smoothing quantification algorithm (Spec ¶¶39-45), integrating the exception into a practical application.
Dependent claim(s) 12-14 and 17-19 recite(s) further steps of classifying, grouping on feature embeddings, obtaining expert labeling, retraining, and extracting and downsampling features, all recited as results to be achieved, which do(es) not integrate the exception into a practical application or amount to significantly more, for the reasons given above.
CLAIM INTERPRETATION
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: the image retrieval network, the region segmentation network, the IVC quantification and RAP estimation network, the classification engine, and the clustering engine, in claim(s) 1-5, 6, 9, and 10.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-4, 11, 12, 16 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al., CN 113506270 A (hereinafter “ZHANG”) in view of Albani et al., “Accuracy of right atrial pressure estimation using a multi-parameter approach derived from inferior vena cava semi-automated edge-tracking echocardiography: a pilot study in patients with cardiovascular disorders hereinafter “ALBANI”), in further view of Caplan et al., “Measurement site of inferior vena cava diameter affects the accuracy with which fluid responsiveness can be predicted in spontaneously breathing patients: a post hoc analysis of two prospective cohorts”(hereinafter “CAPLAN”), and in further view of Zamzmi et al., “Open World Active Learning for Echocardiography View Classification” (hereinafter “ZAMZMI”), or in the alternative to Zamzmi, further in view of Shu et al., “Unseen Class Discovery in Open-World Classification” (hereinafter “SHU”) and, as to the Shu alternative only, in further view of Bendale et al., “Towards Open World Recognition” (hereinafter “BENDALE”).
Claims 1, 11, and 16.
Zhang teaches an Artificial Intelligence (Al) system for estimating inferior vena cava (IVC) collapsibility and
an image retrieval network, the image retrieval network configured to receive images from the echocardiography study and performs a quality assessment of the images to generate selected images (ZHANG: “The measurement algorithm uses a multi-task deep learning network to achieve inferior vena cava segmentation and image quality evaluation (S704). The image quality is divided into three types: high, medium, and low.” (p. 9) and “The preset condition is that the segmentation output 103 is used to determine the diameter of the inferior vena cava when the quality is classified as high or medium quality.” (p. 6) Only frames meeting the quality condition are passed forward, i.e. selected images);
a region segmentation network, the region segmentation network configured to localize and segment the selected images to obtain IVC regions in the selected images (ZHANG: “the multi-task deep neural network 200 uses the improved algorithm network structure of the UNET network segmentation algorithm, and the multi-task deep neural network 200 uses the medical image 102 as input to generate the mask image data and image quality of the inferior vena cava. The category label is output. The multi-task deep neural network 200 uses 4 times of downsampling and 4 times of upsampling to obtain the segmentation output.” (p. 4));
an IVC quantification and RAP estimation network, the IVC quantification and RAP estimation network configured to perform a distance calculation to find a diameter in the IVC regions at (ZHANG: "the diameter measuring unit 32 determines the diameter of the inferior vena
cava based on the intersection between the upper boundary (corresponding to the blood
vessel wall) and the lower boundary ( corresponding to another blood vessel wall) of the
detection line and the contour of the segmented inferior vena cava" and "during the
measurement, a time series sequence of medical images is formed, a plurality of inferior vena cava diameters are determined, and the plurality of inferior vena cava diameters are
displayed in a waveform according to the time sequence". The diameter is found over
successive frames, i.e. at different temporal points.); and
Claim 16 recites the same steps as instructions stored on a non-transitory computer readable medium (ZHANG: “In a third aspect, the present disclosure provides a computer-readable storage medium with instructions stored on the computer-readable storage medium, which
implement the steps of any one of the foregoing methods when executed by one or more
processors.” (p. 2))
Zhang does not specifically teach estimating right atrial pressure. However,
Albani in the same field of endeavor teaches deriving right atrial pressure (ALBANI: “The echocardiographic estimation of right atrial pressure (RAP) is based on the size and inspiratory collapse of the inferior vena cava (IVC)." (p. 1213, Abstract). RAP is obtained from two values Zhang already computes “minimum and maximum diameters of the inferior vena cava and displayed on the interface in real time” (p. 8).)
Zhang and Albani do not specifically teach finding the diameter at different spatial (different sites) points. However, Caplan in the same field of endeavor teaches measuring the IVC diameter at a plurality of sites along the vessel (CAPLAN: “Measurement of IVC diameters were carried out at five sites: at the IVC-right atrium junction (site 0), then at 1 (site 1), 3 (site
3), 4 (site 4) and 5 cm (site 5) caudal to the IVC-atrial junction” (Fig. 1 caption, p. 3), and
“Measuring inferior vena cava diameters 4 cm caudal to the right atrium predicted fluid
responsiveness with the best accuracy.” (p. 1, Abstract).)
Zhang, Albani and Caplan do not specifically teach the open-world active learning system. However, Zamzmi in the same field of endeavor teaches this system applied to
echocardiography view classification (ZAMZMI: “Our proposed framework has two main
stages: a classification stage and a clustering stage.” (p. 305, § 3.2).). In the alternative, Shu in the same field of endeavor teaches an open-world classifier that rejects inputs belonging to classes not seen in training and then clusters those rejected inputs to discover the classes hidden in them (SHU: “This paper concerns open-world classification, where the classifier not only needs to classify test examples into seen classes that have appeared in training but also reject examples from unseen or novel classes that have not appeared in training.” (p. 1, Abstract).)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the multitask IVC network of Zhang to convert the measured diameters into a right atrial pressure value as taught by Albani, to measure at different sites along the vessel as taught by Caplan, and to add the open world active learning stage of Zamzmi/Shu, because converting the measured diameters and collapse index into a right atrial pressure value would provide the pressure estimate Zhang leaves unproduced, and Zamzmi's cluster-and-label update or Shu’s open-world classifier that rejects examples from unseen classes and clusters the rejected examples to discover the hidden classes in them would let Zhang's quality classifier handle a view absent from training instead of failing silently. This is applying a known technique to a known device ready for improvement to yield predictable results (MPEP 2143(D)).
Claim 2.
Zhang, Albani, Caplan, Zamzmi/Shu teaches wherein the classification engine is configured to detect known views in the selected images and unknown views in the selected images (ZAMZMI: “In the classification stage, an open world view classifier is trained to classify known echocardiography views and detect unknown views.” (§ 3.2)). In the alternative, Shu teaches the same limitation (SHU: “build a model f(x) that can classify each test example x to one of the m seen classes in S or reject it as unseen” (p. 2) and “OCN is used for open classification (traditional classification with rejection capability), which can produce rejected examples when tested on both seen and unseen class examples.” (p. 3). The rejected examples are the recited unknown views.).
Claim 3.
Zhang, Albani, Caplan, Zamzmi/Shu and Bendale teaches wherein the clustering engine is configured to group the unknown views into one or more clusters to use to update the classification engine (ZAMZMI: “In the clustering stage, similar unknown images are grouped into clusters to be labeled by a human expert before passing them back to the classification stage for model update.” (§ 3.2).).
In the alternative, Shu teaches the clustering of the rejected views (SHU: “Hierarchical
clustering clusters the rejected examples from OCN using PCN as the distance function. It
gives the number of hidden clusters or classes embedded in the rejected examples.” (p. 3).)
Shu does not itself perform the model update, stating (SHU: “Our work does not focus on
incremental learning.” (p. 2).) Shu identifies the missing step and names where it is found
(SHU: "an open classification system should be able to (1) assign each incoming/test
example to one of the seen classes (appeared in training) and reject those examples from
hidden unseen classes (not appeared in training), (2) discover hidden unseen classes in the
rejected examples, and (3) learn the new classes incrementally. Since many existing
methods exist for (1) and (3) (Scheirer et al., 2013; 2014; Bendale & Boult, 2015; Fei & Liu,
2016; Bendale & Boult, 2016), we will not study them in this paper.” (p. 1).) Bendale
supplies the update (BENDALE: “for open world recognition, we propose the system
explicitly label novel inputs as unknown and then incrementally incorporate them into the
classifier.” (p. 1893).).
Claims 4, 12, and 17.
Zhang, Albani, Caplan, Zamzmi/Shu and Bendale teaches wherein the open-world active learning system is configured to: classify the known views into known classes and the unknown views into an unknown class; group the unknown views of the unknown class based on feature embeddings of the selected images determined by the image retrieval network (ZAMZMI: “the embedded features of the unknown images, which are extracted by the autoencoder (see Figure 2), are used by K-medoids to generate k clusters of unknown classes” (§ 3.2.2));
obtain expert labeling of the unknown views of the unknown class;
add the expert labeling to the unknown views of the unknown class to determine an updated class (ZAMZMI: “a certified echocardiologist labeled each cluster group of unknown images instead of labeling all the unknown images” (§ 3.2.2)); and
retrain the classification engine of the open-world active learning system to know the updated class (ZAMZMI: “the newly labeled images (previously unknown) were sent back to the classification stage and used to update the open world classifier” (§ 3.2.2).).
In the alternative, Shu teaches grouping on feature embeddings (SHU: “Auto-encoder is used to learn representations from unlabeled examples.” (p. 3). The learned representations are the recited feature embeddings.) and Bendale teaches the expert labeling and the retraining
(BENDALE: “for open world recognition, we propose the system explicitly label novel inputs as unknown and then incrementally incorporate them into the classifier.” (p. 1893).).
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG in view of ALBANI in view of CAPLAN in view of ZAMZMI, or in the alternative to Zamzmi, in view of SHU and BENDALE, further in view of Jaderberg et al., “Spatial Transformer Networks” (hereinafter “JADERBERG”).
Claim 5.
Zhang, Albani, Caplan, Zamzmi/Shu and Bendale do not teach wherein the region segmentation network comprises a spatial transformation network to focus on specific regions the selected images.
However, Jaderberg in the same field of endeavor teaches this (JADERBERG: “the
spatial transformer module is a dynamic mechanism that can actively spatially transform
an image (or a feature map) by producing an appropriate transformation” and "spatial
transformers can be trained with standard back-propagation" (p. 1).).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the spatial transformer of Jaderberg into the region segmentation network of Zhang, because focusing the network on the vessel region before segmentation lets the segmentation network spend capacity on the vessel rather than on surrounding tissue. This is applying a known technique to a known device ready for
improvement to yield predictable results (MPEP 2143(D)).
Claim(s) 6-8, 13-15, and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG in view of ALBANI in view of CAPLAN in view of ZAMZMI, or in the alternative to Zamzmi, in view of SHU and BENDALE, in view of JADERBERG, further in view of Yu et al., “BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation” (hereinafter “YU”) and Anwer et al., “Binary Patterns Encoded Convolutional Neural Networks for Texture Recognition and Remote Sensing Scene Classification” (hereinafter “ANWER”).
Claims 6, 13, and 18.
Zhang, Albani, Caplan, Zamzmi/Shu and Bendale, and Jaderberg teach the system according to claim 5, but do not teach wherein the region segmentation network further comprises: a spatial pathway to extract low-level details of the specific regions of the selected images; a handcrafted pathway to extract rich texture features of the specific regions of the selected images; and a context pathway for downsampling of a feature map to obtain a receptive field.
Zhang uses a single encoder-decoder. However, YU teaches the spatial pathway (YU: “The spatial information captured by the Spatial Path encodes mostly rich detail information … In other words, the output feature of Spatial Path is low level” (p. 7).) and the context pathway (YU: “The Context Path utilizes lightweight model and global average pooling [5,6,21] to provide large receptive field. In this work, the lightweight model, like Xception [8], can downsample the feature map fast to obtain large receptive field” (p. 6).).
Zhang and Yu do not teach the handcrafted pathway. However, Anwer teaches a parallel stream carrying handcrafted texture features (ANWER: "we propose deep models, which we call TEX-Nets, by designing a two-stream deep architecture where texture coded mapped images are used as the second stream and fuse it with the normal RGB image stream. To obtain the texture coded mapped images, we first extract LBP based codes from an image." (p. 4).).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to restructure the region segmentation network of Zhang into a spatial pathway, a handcrafted pathway, and a content pathway as taught by Yu and Anwer, because Yu and Anwer both split detail from context into separate concurrent pathways, and Anwer shows a handcrafted texture stream suits speckle ultrasound, so a skilled artisan would separate Zhang’s singe stack accordingly. This is combining prior art elements according to known methods to yield predictable results (MPEP 2143(A))
Claims 7, 14, and 19.
Zhang, Albani, Caplan, Zamzmi/Shu and Bendale, Jaderberg, Yu, and Anwer teach wherein outputs of the spatial pathway, the handcrafted pathway, and the context pathway are combined using a path fusion to obtain a weighted feature vector (YU: “Given the different level of the features, we first concatenate the output features of Spatial Path and Context Path. And then we utilize the batch normalization [15] to balance the scales of the features. Next, we pool the concatenated feature to a feature vector and compute a weight vector, like SENet [13]. This weight vector can re-weight the features, which amounts to feature selection and combination.” (p. 7).)).
Claims 8, 15, and 20.
Zhang, Albani, Caplan, Zamzmi/Shu and Bendale, Jaderberg, Yu, and Anwer teaches wherein the path fusion comprises: concatenating the outputs of the spatial pathway, the handcrafted pathway, and the context pathway to obtain concatenated features; combining the concatenated features into a feature vectors ending the feature vector to a global pooling (Yu teaches each recited operation in the recited order: the concatenation, batch normalization, pooling and re-weighting in the prose at p. 7 quoted above.); performing a 1X1 convolution on the feature vector after the global pooling;
performing a Rectified Linear Unit (ReLU) activation on the feature vector after the 1X1 convolution; performing a second 1X1 convolution on the feature vector after the ReLU activation; applying a Sigmoid function to generate a weight vector; and re-weighting the feature vector based on the weight vector to obtain the weighted feature vector (YU: Figure 2(c) at p. 5, shows concatenate, conv+bn+relu, global pool, 1x1 conv, relu, 1x1 conv, sigmoid, mul and add in that sequence).
PNG
media_image1.png
188
413
media_image1.png
Greyscale
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG in view of ALBANI in view of CAPLAN in view of ZAMZMI, or in the alternative to Zamzmi, in view of SHU and BENDALE, further in view of El Harouni et al., US 2020/0051238 A1, (hereinafter “EL HAROUI”).
Claim 9.
Zhang, Albani, Caplan, Zamzmi/Shu teaches the system according to claim 1, wherein the image retrieval network comprises: a shared encoder; a first task-specific head; and a second task-specific head, wherein the first task-specific head and the second task-specific head (ZHANG: “the multi-task deep neural network 200 includes: an encoder sub-network 201, a decoder sub-network 202, a link block 203, and a classification sub-network 204” and “Thus, the segmentation and classification processes share at least part of the network layer” (p. 4).).
Zhang does not state how many layers each branch contains. However, El Harouni the
same field of endeavor teaches a shared encoder carrying two task-specific heads (EL HAROUNI: “The network architecture is based on the U-net architecture with two output
heads, one for segmentation and the other for classification.” and “The network consists of
a concatenating path and an expanding path, the classification and segmentation outputs
are at the end of the concatenating and expanding paths, respectively.” (¶ 71).) and
teaches the layers of each head (EL HAROUNI: “For the classification task, the network
includes an added global average pooling (GAP) layer at the end of the concatenating path,
followed by a convolution layer with seven filters corresponding to the seven classes and finally a softmax layer to produce the class probabilities.” (¶ 72). Global average
pooling, convolution and softmax are three task-specific layers.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to provide each task-specific head of the shared-encoder network of Zhang with the layer structure taught by El Harouni, because each branch needs its own defined stack of layers to process the shared encoder features separately for segmentation and quality classification. This is applying a known technique to a known device ready for improvement to yield predictable results (MPEP 2143(D)).
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG in view of ALBANI in view of CAPLAN in view of ZAMZMI, or in the alternative to Zamzmi, in view of SHU and BENDALE, further in view Sandler et al., “MobileNetV2: Inverted Residuals
and Linear Bottlenecks” (hereinafter “SANDLER”).
Claim 10.
Zhang, Albani, Caplan, Zamzmi/Shu teaches the system according to claim 1, wherein the region segmentation network is a lightweight region segmentation network configured to achieve an (ZHANG: “each of the foregoing encoder blocks includes one or more convolutional neural network layers, and the convolutional neural network layer is a separable convolutional neural network layer” (p. 2).).
Zhang does not recite a frame rate. However, Sandler teaches the lightweight architecture applicant itself builds its network from (SANDLER: “In this paper we
describe a new mobile architecture, MobileNetV2, that improves the state of the art
performance of mobile models on multiple tasks and benchmarks as well as across a
spectrum of different model sizes.” (p. 1).) See also the specification ¶ 23, which builds the image retrieval network from “five inverted residual bottleneck (see MobileNetV2-s residual
bottleneck [8]) blocks and a final pooling layer.” The recited frame rate is a result of network
depth and hardware and is obtained through routine optimization. See In re Aller, 220 F.2d 454
(CCPA 1955); MPEP 2144.05(II).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to build the lightweight region segmentation network of Zhang from the MobileNetsV2 architecture of Sandler, because both use depthwise separable convolutions for efficient mobile inference and applicant’s specification combines them in the same retrieval network. This is a simple substitution of one known element for another to obtain predictable results (MPEP 2143(B)).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ross Varndell whose telephone number is (571)270-1922. The examiner can normally be reached M-F, 9-5 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’Neal Mistry can be reached at (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Ross Varndell/Primary Examiner, Art Unit 2674