DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “loss application module” in claim 10.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-19 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites “a similarity” in line 13. It is unclear what this similarity refers to and what is being compared to determine the similarity.
Claim 6 recites the limitation "check whether data is augmented " in line 2. It is unclear what “data” is being referred to, and thus the claim is indefinite.
Claim 6 line 3 recites “perform encoding for whether data is augmented”. It is unclear what is being encoded and whether the encoding is to determine whether the data has been augmented or on the data that is augmented, thus leading to indefiniteness.
Claim 6 line 3 recites “perform encoding for whether data”. The term “data” is unclear whether the it is referring to the data of the previous line of is new data, thus leading to indefiniteness.
Claim 6 line 7 recites “a text domain” which is previously introduce in claim 3 which claim 6 dependents on thus creating a lack of clarity of whether this refers to a new text domain or the same text domain, thus leading to indefiniteness.
Claim 11 recites “a similarity” in line 8. It is unclear what this similarity refers to and what is being compared to determine the similarity.
Claim 14 recites the limitation "check whether data is augmented " in line 2. It is unclear what “data” is being referred to, and thus the claim is indefinite.
Claim 14 line 3 recites “perform encoding for whether data is augmented”. It is unclear what is being encoded and whether the encoding is to determine whether the data has been augmented or on the data that is augmented, thus leading to indefiniteness.
Claim 14 line 3 recites “perform encoding for whether data”. The term “data” is unclear whether the it is referring to the data of the previous line of is new data, thus leading to indefiniteness.
Claim 14 line 7 recites “a text domain” which is previously introduce in claim 3 which claim 6 dependents on thus creating a lack of clarity of whether this refers to a new text domain or the same text domain, thus leading to indefiniteness.
Claim 18 recites “a similarity” in line 10. It is unclear what this similarity refers to and what is being compared to determine the similarity.
Claim 2-5, 7-10, 12-14, 15-17, and 19 are rejected due to their dependency on the rejected base claim.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim 1, 11, and 18 is rejected under 35 U.S.C. 102(a)(1) as being anticipated by Yuan “Multimodal Contrastive Training for Visual Representation Learning” (hereinafter “Yuan”).
Regarding claim 1, Yuan teaches an electronic device including a pretraining module, a loss application module, a score application module, the electronic device comprising (see section 3 and 3.1, a system with training schemes that includes losses with similarity scores. The initial model is pretrained (see section 4.2)):
a memory (see section 4.1 Implementation details, the use of a GPU which includes memory); and
a processor configured to provide a pretraining unified framework based on contrastive text image stored in the memory by controlling operations of the pretraining module, the loss application module, and the score application module (see section 3 Method and Figure 2, a multi-modal contrastive training framework that includes inter-modal contrastive learning (between image and caption). The method system includes pretrained model, contrastive learning with losses, and the use of similarity scores [section 3.1]. The use of the methods on a GPU which includes memory, section 4.1)
wherein the processor is configured to (see section 4.1, the use of a GPU):
perform pretraining on a data set including at least one of text and images corresponding to a data set domain input through the pretraining module (see section 4.1 pretraining dataset, the training of the model on the image-caption-tag tubles of a dataset);
apply a loss to a plurality of positive samples in the pretrained data set through the loss application module (see section 3.1-3.2, the intra-modality and inter-modality contrastive learning using equations 3, 8, 11, and 14 to calculate losses using
k
i
i
+
,
k
c
c
+
,
k
c
i
+
, and
k
i
c
+
to denotes that a positive same is being formed with the features from the image and text inputs. All the losses calculated are used for a final loss in equation 15. The features denote in the loss equations are based on the multi-modal dataset which is broken into vectors of Images, captions, and tags to be embedded into features. The multi-modal dataset which is created based on the COCO and Stock pretraining datasets, section 4.1); and
apply a score for embedding pretrained data sets from a plurality of domains in the same space based on a similarity through the score application module (see section 3 paragraph 2 and section 3.1-3.2, the calculation of similarity scores between example pairs of the image and text inputs [the example pairs are embedded into features
k
i
i
j
,
k
c
c
j
,
k
c
i
j
,
k
i
c
j
,
q
i
i
j
,
q
c
c
j
,
q
c
i
j
, and
q
i
c
j
]. The similarity score is computed by the dot product with embedded features in equations 3, 8, 11, and 14 . The visual and textual features can be embedded into a common space).
Claims 11 and 18 are analogous to claim 1, and thus similarly analyzed and rejected as claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 2-4, and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Yuan in view of You et al “Graph Contrastive Learning Automated” (included in IDS)(hereinafter “You”).
Regarding claim 2, Yuan teaches the electronic device of claim 1.
Yuan teaches the processor is configured to perform pretraining on the data set domain through the pretraining module based on an augmentation-agnostic image encoder and an (see Figure 2 and section 3-3.1, the training includes use of encoders for augmented image and captions [augmentation-agnostic image encoders] and a 2-layer MLP head [projection head]).
Yuan does not teach an augmentation-aware projection head.
You teach an augmentation-aware projection head (see section 1 page 2, an augmentation -aware projection head for contrastive learning).
Yuan and You are analogous art because they are from the same field of endeavor of methods of contrastive learning for augmented data in an automated setting.
Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Yuan to use an augmentation-aware projection head as taught by Yuan. The motivation for doing so would have been to limit the complicated augmentations distorting the original data distributions (You, Section 1).
Regarding claim 3, Yuan and You teach the electronic device of claim 2.
Yuan teaches the processor is configured to perform pretraining to which data augmentation is applied on a text domain, an image domain, and a text-image composite domain through the pretraining module (see section 3.1- 3.2 and Figure 2, the augmentation of images, text during training. See Figure 2 and section 3.2, the use of image augmentation and caption augmentation together for intermodal contrastive learning. The image and text are both augmented and then used together as a pair for contrastive learning, thus creating an augmented text-image domain that is used in the training scheme) , and
wherein the image domain includes a basic image domain, a first-stage augmentation image domain, and a second-stage augmentation image domain, which are embedded in the same space (see section 3.1, images are augmented to produce different variants of the same input image, thus there is the original image [basic image domain], and augmentation variants or examples [first and second stage segmentation image domains]).
Regarding claim 4, Yuan and You teach the electronic device of claim 3.
Yuan teaches the first-stage augmentation image domain and the second-stage augmented image domain are generated by applying different augmentation techniques (see section 3.1, images are augmented to produce different variants of the same input image. See section 4.1 Implementation Detail, the augmentation is performed according to the data augmentation scheme which includes using image cropping, color jittering, horizontal flipping, greyscale conversion and gaussian blurring. Thus, different techniques can be used to produce the different variants)
wherein the augmentation techniques include at least one augmentation technique among brightness adjustment, contrast adjustment, rotation, scaling, and color distortion (See section 4.1 Implementation Detail, the augmentation is performed according to the data augmentation scheme which includes using image cropping, color jittering [color jittering includes adjustment of brightness, contrast, and hue], horizontal flipping, greyscale conversion and gaussian blurring).
Claims 12-13 are analogous to claim 2 and3, respectively, and thus similarly analyzed and rejected as claim 2, and 3.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Yuan in view of You in view of Chen US20210327029 (hereinafter “Chen”).
Regarding claim 5, Yuan and You teach the electronic device of claim 4.
wherein the processor is configured to (see section 4.1, the use of a GPU):
perform pretraining such that data of the first-stage augmentation image domain is generated by image augmenting data of the basic image domain through a weak augmentation technique including brightness adjustment and contrast adjustment (see section 3.1 and 4.1, images [basic image] are augmented to produce different variants of the same input image. The augmentation is performed according to the data augmentation scheme which includes color jittering [color jittering includes adjustment of brightness, contrast, and hue]. The use of augmentation scheme color jittering to produce an augmentation variant of the image is interpreted as the generated the first-stage augmentation image domain using adjustments of brightness and contrast), and
perform pretraining such that data of the second-stage augmentation image domain is generated by image augmenting the data of the basic image domain (see section 3.1, images are augmented to produce different variants of the same input image, thus there is the original image [basic image domain], and augmentation variants or examples [first and second stage segmentation image domains]).
You nor Yuan teach image augmenting the data of the basic image domain through a strong augmentation technique including rotation, scaling, and color distortion.
Chen teaches image augmenting the data of the basic image domain through a strong augmentation technique including rotation, scaling, and color distortion (see paragraph 0056, data augmentation [see 0038 the data is an input image, interpreted as the basic image domain] includes geometric transformation of data such as resizing [scaling] and rotation and augmentation involving appearance transformation including color distortion).
Chen, Yuan and You are analogous art because they are from the same field of endeavor of methods of contrastive learning for augmented data in an automated setting.
Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Yuan and You to use augmentation techniques of scaling, rotation, and color distortion as taught by Chen. The motivation for doing so would have been to benefit the contrastive learning, understand the importance of augmentation composition, and adjust the original images for differences (Chen, 0056-0062).
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Yuan in view of You in view of Lu US20230162481 (hereinafter “Lu”).
Regarding claim 19, Yuan teaches the chipset of claim 18.
Yuan does not teach the at least one integrated circuit comprises at least one of Programmable Gate Array (FPGS) and Application-Specific Integrated Circuit (ASIC).
Lu teaches the at least one integrated circuit comprises at least one of Programmable Gate Array (FPGS) and Application-Specific Integrated Circuit (ASIC) (see paragraph 0073, aspects of the logic machine and storage machine may be integrated together into hardware components which may include FPGAs and ASICs).
Yuan and Lu are analogous art because they are from the same field of endeavor of methods of contrastive learning for augmented data comprising image-text pairs in an automated setting.
Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Yuan to use FPGS or ASIC as taught by Lu. The motivation for doing so would have been to integrate machines together into hardware logic components (Lu, 0073).
Conclusion
Claims 6-10 and 14-17 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see the attached 892 notice of reference cited.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to EMILY R. HAUK whose telephone number is (571)272-5966. The examiner can normally be reached M-F 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/EMILY ROSE HAUK/
Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669