DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Remarks
The Office Action has been made issued in response to Response to Election/Restriction Requirement filed January 15th, 2026. Claims 1-14 are pending.
Election/Restrictions
Claims 15-20 of group II are withdrawn from further consideration pursuant to 37 CFR 1.142(b) as being drawn to a nonelected groups, respectively, there being no allowable generic or linking claim. Election was made without traverse in the reply filed on June 15th, 2026. Accordingly, claims 1-14 are being based on the elected group I by the applicants.
Preliminary Amendment
The preliminary amendment filed on October 14th, 2024 has been acknowledged and entered.
Claim Objections
Claims 1 and 9 are objected to because of the following informalities:
Claim 1, line 8, “so as to output candidate” should be read as “” to follow proper claim language. Appropriate correction is required.
Claim 1, line 5, the reference “via operation of the one or more annotation specialist models” should be read as “” to follow proper claim language formality. Appropriate correction is required.
Claim 1, line 7, the reference “via operation of a data filtering and enhancement module” should be read as “”
Claim 9, line 8, “so as to output candidate” should be read as “” to follow proper claim language. Appropriate correction is required.
Claim Interpretation
6. The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitation(s) that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function.
Claim(s) 1 and 9, recite(s) limitation(s) that use words like “means” (or “step”) or similar terms with functional language that do invoke 35 U.S.C. 112(f):
Claim 1; recites the limitation, “receiving, at one or more annotation specialist models, a plurality of images to be annotated; via operation of the one or more annotation specialist models, generating pre-filtered annotations for the plurality of images,” [Lines 3-6].
Claim 1; recites the limitation, “via operation of a data filtering and enhancement module, filtering the pre-filtered annotations….,” [Line 7].
Claim 9; recites the limitation, “one or more annotation specialist models configured to: receive a plurality of images to be annotated; and generate pre-filtered annotations for the plurality of images,” [Lines 3-5].
Claim 9; recites the limitation, “…enhancement module configured to: filter the pre-filtered annotations in accordance…,” [Lines 6-8].
Claim 9; recites the limitation, “an iterative data refinement model configured to: iteratively train the multi-task computer….,” [Lines 9-11].
Claim 9; recites the limitation, “a final annotation module configured to store a candidate annotation…,” [Lines 12-14].
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 1 and 9;
(i) “one or more annotation specialist models” which has no sufficient written support of a structure, material and/or act for the recited feature to perform the recited function, the closest disclosure can be found in the specification filed on January 30th, 2024, in Par. [0035] which discloses “annotation specialist models includes trained optical character recognition OCR application programming interface, trained caption model, trained grounding model, trained object/proposal determination model, and trained segmentation model”, however, doesn’t disclose specific algorithm and/or structure with adequate details for understanding that the recited can perform the recited function. thus have no sufficient structure, material and/or act.
(ii) “enhancement module” which has sufficient written support of a structure, material and/or act for the recited feature to perform the recited function, wherein the data filtering and enhancement module is disclosed, in the specification filed on January 30th, 2024, in Par. [0039] which discloses “data filtering and enhancement module comprises both a text filter and enhancement module and a region filtering model. Text filter and enhancement module comprises large multi-modal annotator, large language model annotator, and text filter. Region filtering module comprises region score model, non-maximum suppression model, blacklist, and previously trained foundation model, thus have sufficient structure or material wherein the data filtering and enhancement module comprises both a text filter and enhancement module and a region filtering model, the text filter and enhancement module comprises large multi-modal annotator, large language model annotator, and text filter, the region filtering module comprises region score model, non-maximum suppression model, blacklist, and previously trained foundation model.
(iii) “an iterative data refinement model” the disclosure, filed on January 30th, 2024, has no written support for a structure, material/act for the recited iterative data refinement model to perform the recited function, the closest support can be found in Par. [00166] wherein it discloses “an interactive data refinement model configured to iteratively train….” which is a nominal mentioning of the feature and restatement of the claim language, thus have no sufficient structure or material/act.).
(iv) “a final annotation module” the disclosure, filed on January 30th, 2024, has no written support for a structure, material/act for the recited iterative data refinement model to perform the recited function, the closest support can be found in Par. [00166] wherein it discloses “a final annotation module configured to store the candidate annotation…” which is a nominal mentioning of the feature and restatement of the claim language, thus have no sufficient structure or material/act.).
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1 and 9 along with their dependent claims 2-8 and 10-14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim 1’s limitations:
Claim 1; recites the limitation, “receiving, at one or more annotation specialist models, a plurality of images to be annotated; via operation of the one or more annotation specialist models, generating pre-filtered annotations for the plurality of images,” [Lines 3-6].
Claim 9’s limitations:
Claim 9; recites the limitation, “one or more annotation specialist models configured to: receive a plurality of images to be annotated; and generate pre-filtered annotations for the plurality of images,” [Lines 3-5].
Claim 9; recites the limitation, “an iterative data refinement model configured to: iteratively train the multi-task computer….,” [Lines 9-11].
Claim 9; recites the limitation, “a final annotation module configured to store a candidate annotation…,” [Lines 12-14].
Claims 1 and 9, each respectively invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The specification is devoid of adequate structure to perform the claimed functions. The specification does not provide sufficient details such that one of the ordinary skill in the art would understand which structure performed(s) the claimed function.
Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph.
Applicant may:
(a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
(b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
(a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1 and 9 along with their dependent claims 1-8 and 10-14 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention. As described above, the disclosure does not provide adequate structure to perform the claimed function in the recited limitation.
Claim 1’s limitations:
Claim 1; recites the limitation, “receiving, at one or more annotation specialist models, a plurality of images to be annotated; via operation of the one or more annotation specialist models, generating pre-filtered annotations for the plurality of images,” [Lines 3-6].
Claim 9’s limitations:
Claim 9; recites the limitation, “one or more annotation specialist models configured to: receive a plurality of images to be annotated; and generate pre-filtered annotations for the plurality of images,” [Lines 3-5].
Claim 9; recites the limitation, “an iterative data refinement model configured to: iteratively train the multi-task computer….,” [Lines 9-11].
Claim 9; recites the limitation, “a final annotation module configured to store a candidate annotation…,” [Lines 12-14].
The specification does not demonstrate that applicant has made an invention that achieves the claimed function because the invention is not described with sufficient detail such that one of ordinary skill in the art can reasonably conclude that the inventor had possession of the claimed invention.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-3 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Maximilian Metzner et. al. (“US 2024/0386637 A1” hereinafter as “Metzner”) in view of Hanoona Abdul Rasheed et. al. (“US 2024/0203085 A1” hereinafter as “Rasheed”) further in view of Hao-Cheng Kao et. al. (“US 2020/0065623 A1” hereinafter as “Kao”) and Guannan Jiang et. al. (“US 2023/0267716 A1” hereinafter as “Jiang”).
Regarding claim 1, Metzner teaches a method for annotating images to create a corpus (Pars. [0059-60] discloses “by repeating steps S2 and S3, large training data image sets may be rapidly generated…for n objects…unique variations in the presence of the objects may be generated…different images may therefore be produced from an image 253 for training” indicating generating a corpus [“large training data image sets with unique variations”]) for training (Par. [0002] discloses “method and systems for providing or for generating and providing training image data for training a function”) a multi-task computer vision machine learning model, comprising (moreover, Pars. [0062-65] discloses “the training data TBD may be used to train a function, for example an object recognition function…training object recognition functions for systems that are used, for example, in the field of autonomous driving”; Par. [0032] discloses “annotation contains information about the size and position of the at least one object in the annotated image and/or a segmentation, i.e., information about all the pixels associated with the object” indicating a multi-task model including labeling, segmenting, annotating, size and position determination [analogous to multi-task object recognition functions computer vision machine learning model]): receiving, at one or more annotation specialist models (Par. [0041] discloses “FIG. 1 depicts a flow chart of a computer-implemented method…depicts possible interim results of the methods steps of the method”, the programmed processors to perform this function is analogous to the recited one or more annotation specialist models as claimed), a plurality of images to be annotated (Par. [0051] discloses “a multiplicity of modified images MAB may be generated, that (for example, together with the annotated image AB) are provided as training image data”; the programmed processors to perform this function is analogous to the recited one or more annotation specialist models as claimed); via operation of the one or more annotation specialist models, generating pre-filtered annotations for the plurality of images (Par. [0052] discloses “a further image HB. In the image HB, a (partially) unpopulated printed circuit board may be seen, on which none of the objects are depicted”; moreover, Par. [0056] discloses “it may be expedient that the area of the background image HB with which the image are defined by the label L2 is overwritten, and the area to be overwritten itself correspond to one another” therefore, the area in HB image is to be overwritten [analogous to pre-filtered], and the image area of the HB image is defined by a label [analogous to annotation] hence, the image HB with its areas being analogous to pre-filtered annotations for the plurality of images as claimed); via operation of a data filtering module (Par. [0041] discloses “FIG. 1 depicts a flow chart of a computer-implemented method…depicts possible interim results of the methods steps of the method”, the programmed processors to perform this function is analogous to the recited a data filtering module as claimed), filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images (Par. [0057] discloses “the modified image MAB has been produced, for example, by such overwriting. Such a modified image MAB is highly realistic and improves the quality and accuracy of a function when it is trained” indicating a modified image by overwriting and modifying of the image to produce highly realistic image, hence is a product of filtering since it removes/modifies/overwrites objects from the image [filtering out objects]; the programmed processors to perform this function is analogous to the recited a data filtering module as claimed); selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image (“or” indicates selection, therefore, only one of these options is the instant scope of the claim, the examiner selects “selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image”, wherein Metzner’s Pars. [0059-60] discloses “by repeating steps S2 and S3, large training data image sets may be rapidly generated…for n objects…unique variations in the presence of the objects may be generated…different images may therefore be produced from an image 253 for training” indicating generating a corpus [“large training data image sets with unique variations”] of training data with label for these associated images), or (2) adding the candidate annotation to its associated image using the one or more annotation specialist models and the data filtering and enhancement module for subsequent iterative annotation and filtering.
However, Metzner does not explicitly teach via operation of an enhancement module, filtering the pre-filtered annotations.
Rasheed teaches via operation of an enhancement module (elements of 112f interpretation, Par. [0004] discloses “object detection is a foundation for high-level tasks” indicating the object detection model here is a foundation model being pre-trained; Par. [0017] discloses “object detection using high-quality object proposals from a pre-trained multi-modal vision transformer” indicating a multi-modal vision transformer being the annotator of an enlarged detector vocabulary; Par. [0038] discloses “the present approach connects the image, region, and language representations to generalize better to novel open-vocabulary objects” indicating a large language model, Par. [0090] discloses “L2 normalization is used on the region and text embeddings before computing the RKD loss and final classification scores” indicating a text normalization/filtering according to a region score model; Rasheed’s Par. [0054] discloses “all the boxes are arranged according to their cls scores. Then, a non-maximum suppression is applied with a threshold…all of the bounding boxes…with another bounding boxes are discarded”, wherein discarding indicating blacklisting), filtering the pre-filtered annotations (Par. [0072] discloses “these matrices are normalized by L2 norm applied row-wise” indicating a normalization/filtering; Par. [0090] discloses “A L2 normalization is used on the region and text embeddings before computing the RKD loss and final classification scores. The L2 normalization is helpful to stabilize the training”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Rasheed of having via operation of an enhancement module, filtering the pre-filtered annotations.
Wherein having Metzner‘s method wherein having via operation of an enhancement module, filtering the pre-filtered annotations.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories. Since both Metzner and Rasheed share the same endeavor of systems that perform object detection and labelling in images. Wherein Metzner’s system improve generating training data with labels more effectively and less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], while Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017].
However, Metzner in view of Rasheed does not explicitly teach filtering the pre-filtered annotations in accordance with predefined noise criteria; and for each of one or more candidate annotations, perform the selectively storing.
Kao teaches filtering the pre-filtered annotations in accordance with predefined noise criteria (Pars. [0054-55] discloses “determining the inter-annotator consistency…if all of the consistencies are higher than a threshold, the labelled results may be determined to be valid for training the AI machine and be fed to the AI machine”, moreover, this consistency filtering is the filtering step in FIG. 10B, moreover, Par. [0051] discloses “since the annotator consistently label the three raw data as the class C, the processor may obtain a high intra-annotator consistency of the annotator after calculation the intra-class correlation coefficient”; therefore, indicating a filtering step of the labels [annotations] according to a threshold [predetermined noise criteria as claimed]); and for each of one or more candidate annotations (wherein Kao’s Par. [0115] discloses “after the annotators finish their labeling operations, the first raw data with the shown bounding regions (i.e., labelled data) may be referred to as a first labelled result and retrieved by the processor. With the first labelled result, the processor may accordingly determine whether the first labelled result is valid for training…based on a plurality of consistencies of the annotators” indicating that for each of the label produced by the annotators), perform the selective storing (wherein Kao’s Par. [0115] discloses “after the annotators finish their labeling operations, the first raw data with the shown bounding regions (i.e., labelled data) may be referred to as a first labelled result and retrieved by the processor. With the first labelled result, the processor may accordingly determine whether the first labelled result is valid for training…based on a plurality of consistencies of the annotators”, only when the consistencies between the annotators in their labelling meet certain criteria then, that label is used for training data).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Matzner in view of Rasheed of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module, filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image, or (2) adding the candidate annotation to its associated image using the one or more annotation specialist models and the data filtering and enhancement module for subsequent iterative annotation and filtering, with the teachings of Kao of filtering the pre-filtered annotations in accordance with predefined noise criteria; and for each of one or more candidate annotations, perform the selective storing.
Wherein having Metzner’s method wherein filtering the pre-filtered annotations in accordance with predefined noise criteria; and for each of one or more candidate annotations, perform a selective storing.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to verify generated labelled data to obtain more accurate and robust training data. Since both Metzner and Kao share the same endeavor of systems that perform training data generation. Wherein Metzner’s system improve generating training data with labels more effectively and with less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], while Kao’s system improves generating of training data by performing data verification on generated labelled data to obtain more accurate and robust training data, see Kao’s Pars. [0003-0004].
However, Metzner in view of Rasheed and Kao does not explicitly teach adding the candidate annotation to its associated image using the one or more annotation specialist models for subsequent iterative annotation and filtering.
Jiang teaches adding the candidate annotation to its associated image using the one or more annotation specialist models (Par. [0047] discloses “In step 102, an acquired original image sample is input to the data annotation model to obtain an automatic annotation result corresponding to the original image sample. In step 103, the automatic annotation result is visually presented together with the original image sample based on a preset annotation strategy” wherein the data annotation model is analogous to the recited annotation specialist model) for subsequent iterative annotation and filtering (Par. [0044] discloses “…sending unannotated samples into the model, a pre-annotated result is obtained…the pre-annotation result is directly presented on an original image to form a result map…the corrected result can be used as a new annotated sample to iteratively optimize the image segmentation model”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed and Kao of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module, filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image, with the teachings of Jiang of wherein having adding the candidate annotation to its associated image using the one or more annotation specialist models for subsequent iterative annotation and filtering
Wherein having Metzner’s method wherein having adding the candidate annotation to its associated image using the one or more annotation specialist models for subsequent iterative annotation and filtering
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to increase efficiency of annotating data used as a training sample set, thereby shortening a time for preparing a training data set. Since both Metzner and Jiang share the same endeavor of systems that perform training data generation. Wherein Metzner’s system improve generating training data with labels more effectively and with less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], while Jiang’s system improves efficiency of annotating data used as a training sample set, thereby shortening a time for preparing a training data set, see Jiang’s Par. [0005].
Regarding claim 2, Metzner in view of Rasheed further in view of Kao and Jiang teaches the method of claim 1.
However, Metzner in view of Rasheed further in view of Jiang does not explicitly teach where the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model.
Kao teaches where the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model (“one or more” indicates a selection, therefore, only one of these options is the instant scope of the claim, the examiner selects “trained object proposal and detection models”, wherein Kao’s Par. [0042] discloses “the annotator may recognize the raw data as an image with a cat” indicating an object determination model, moreover, Par. [0060] discloses “certain annotator back to be trained again” indicating the annotators are trained models).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed further in view of Kao and Jiang of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image, or (2) adding the candidate annotation to its associated image using the one or more annotation specialist models and the data filtering and enhancement module for subsequent iterative annotation and filtering, with the teachings of Kao of where the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model.
Wherein having Matzner’s method wherein having the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to verify generated labelled data to obtain more accurate and robust training data. Since both Metzner and Kao share the same endeavor of systems that perform training data generation. Wherein Metzner’s system improve generating training data with labels more effectively and less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], and Kao’s system improves generating of training data by verifying generated labelled data to obtain more accurate and robust training data, see Kao’s Pars. [0003-0004].
Regarding claim 3, Metzner in view of Rasheed further in view of Kao and Jiang teaches the method of claim 1.
However, Metzner in view of Kao and Jiang does not explicitly teach where the filtering of the pre-filtered annotations comprises filtering protocols on text data and region data.
Rasheed teaches where the filtering of the pre-filtered annotations comprises filtering protocols on text data and region data (Par. [0090] discloses “L2 normalization is used on the region and text embeddings” indicating a text normalization or filtering; moreover, Par. [0086] discloses “a dataset for large vocabulary instance segmentation…for pseudo-labeling process” indicates a model for object detection based on segmentation and object pseudo-labeling which is analogous to the model of Metzner, hence it’s obvious to modify Metzner’s object detection model to further include region and text embedding segmentation and normalization for labeling task of images).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed further in view of Kao and Jiang of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Rasheed of where the filtering of the pre-filtered annotations comprises filtering protocols on text data and region data.
Wherein having Matzner’s method wherein having the filtering of the pre-filtered annotations comprises filtering protocols on text data and region data.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories. Since both Metzner and Rasheed share the same endeavor of systems that perform object detection and labelling in images. Wherein Metzner’s system improve generating training data with labels more effectively and less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], and Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017].
Regarding claim 8, Metzner in view of Rasheed further in view of Kao and Jiang teaches the method of claim 1, Metzner teaches further comprising: employing the trained multi-task computer vision machine learning model (Par. [0032] discloses “annotation contains information about the size and position of the at least one object in the annotated image and/or a segmentation, i.e., information about all the pixels associated with the object” indicating a multi-task model including labeling, segmenting, annotating, size and position determination; moreover, the function is to receive images and recognize objects and label) to receive one or more images (Pars. [0062-65] discloses “the training data TBD may be used to train a function, for example an object recognition function…training object recognition functions for systems that are used, for example, in the field of autonomous driving” indicating generating of training data for multi functions [analogous to multi-task object recognition functions computer vision machine learning model]) and to iteratively annotate each of the received images (Par. [0059] discloses “repeating steps S2 and S3, large training data image sets may be rapidly generated” indicating the generation of the training data to train the trained object recognition model includes repeating steps of annotating/labeling, therefore, the employment of the training of the object recognition model include iteratively annotate each of the received images as claimed).
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Maximilian Metzner et. al. (“US 2024/0386637 A1” hereinafter as “Metzner”) in view of Hanoona Abdul Rasheed et. al. (“US 2024/0203085 A1” hereinafter as “Rasheed”) and Hao-Cheng Kao et. al. (“US 2020/0065623 A1” hereinafter as “Kao”) and Guannan Jiang et. al. (“US 2023/0267716 A1” hereinafter as “Jiang”) and Adnan Masood et. al. (“US 2025/0124371 A1” hereinafter as “Masood”).
Regarding claim 4, Metzner in view of Rasheed further in view of Kao and Jiang teaches the method of claim 3.
However, Metzner in view of Rasheed further in view of Kao and Jiang does not explicitly teach wherein the filtering protocol on the text data includes filtering out texts containing excess objects.
Masood teaches wherein the filtering protocol on the text data includes filtering out texts containing excess objects (Par. [0037-38] discloses “for example, the Euclidean distance between the word embeddings of “king” and “queen” can be calculated as follows…the Euclidean distance can be normalized by dividing it by the maximum possible distance in the embedding space” and Par. [0044] discloses “the prominence index is resilient to outliers as an equidistant measure between skills and job description title, and using the harmonic mean of the normalized Euclidean distance…this combined measure is more resilient to outliers, as it down-weights the effect of large differences in either of the distance measures” indicating the normalization of text embedding [analogous to the text embedding normalization of Rasheed] includes down-weighting [filtering out] the effect of outliers [analogous to excess objects]).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed further in view of Kao and Jiang of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Masood of wherein the filtering protocol on the text data includes filtering out texts containing excess objects.
Wherein having Matzner’s method having wherein the filtering protocol on the text data includes filtering out texts containing excess objects.
The motivation behind the modification would have been to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories and further to improve computer technology by improving information extraction in knowledge base representations that may be missed using other information extraction methods. Since both Rasheed and Masood share the same endeavor of systems that perform text embedding normalization. Wherein Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017], and Masood’s system improves computer technology by improving information extraction in knowledge base representations that may be missed using other information extraction methods, see Masood’s Par. [0013].
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Maximilian Metzner et. al. (“US 2024/0386637 A1” hereinafter as “Metzner”) in view of Hanoona Abdul Rasheed et. al. (“US 2024/0203085 A1” hereinafter as “Rasheed”) further in view of Hao-Cheng Kao et. al. (“US 2020/0065623 A1” hereinafter as “Kao”) and Guannan Jiang et. al. (“US 2023/0267716 A1” hereinafter as “Jiang”) and Scott C. Evans et. al. (“US 2004/0257988 A1” hereinafter as “Evans”).
Regarding claim 5, Metzner in view of Rasheed further in view of Kao and Jiang teaches the method of claim 3.
However, Metzner in view of Rasheed further in view of Kao and Jiang does not explicitly teach wherein the filtering protocol on the text data includes retaining texts with a minimum action and object complexity.
Evans teaches wherein the filtering protocol on the text data includes retaining texts with a minimum action and object complexity (Par. [0047] discloses “since the minimum number of phrases in the inputted string is already known…given string is known, the number of phrases in the LZ78 partition can be normalized based on the estimate of the number of phrases…for the normalized complexity estimate” indicating a text normalization process [analogous to text normalization of Rasheed] which includes inputting string of minimum number of phrases [retaining text with a minimum action as claimed] for the normalized complexity is analogous to the recited object complexity as claimed).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed further in view of Kao and Jiang of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Evans of wherein the filtering protocol on the text data includes retaining texts with a minimum action and object complexity.
Wherein having Matzner’s method having wherein the filtering protocol on the text data includes retaining texts with a minimum action and object complexity.
The motivation behind the modification would have been to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories and further to transmit complex data in normalized complexity to allow for efficient transmission of data. Since both Rasheed and Evans share the same endeavor of systems that perform text embedding normalization. Wherein Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017], while Evans’ system improves transmitting complex data in normalized complexity to allow for efficient transmission of data, see Evans’ Abstract.
Claims 6-7 are rejected under 35 U.S.C. 103 as being unpatentable over Maximilian Metzner et. al. (“US 2024/0386637 A1” hereinafter as “Metzner”) in view of Hanoona Abdul Rasheed et. al. (“US 2024/0203085 A1” hereinafter as “Rasheed”) further in view of Hao-Cheng Kao et. al. (“US 2020/0065623 A1” hereinafter as “Kao”) and Guannan Jiang et. al. (“US 2023/0267716 A1” hereinafter as “Jiang”) and Hongwei Zhu et. al. (“US 2013/0176430 A1” hereinafter as “Zhu”).
Regarding claim 6, Metzner in view of Rasheed further in view of Kao and Jiang teaches the method of claim 3.
However, Metzner in view of Rasheed further in view of Kao and Jiang does not explicitly teach wherein the filtering protocol on the region data includes removing noisy boxes under a confidence score threshold.
Zhu teaches wherein the filtering protocol on the region data includes removing noisy boxes under a confidence score threshold (Par. [0032] discloses “determine whether…excluded from consideration as an object…ignore blobs of insignificant size (e.g., below a threshold number of pixels” indicating a region filtering wherein to remove or ignore boxes/blobs below a certain threshold of pixels of size or analogous to confidence score threshold as claimed).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed further in view of Kao and Jiang of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Zhu of wherein the filtering protocol on the region data includes removing noisy boxes under a confidence score threshold.
Wherein having Matzner’s method having wherein the filtering protocol on the region data includes removing noisy boxes under a confidence score threshold.
The motivation behind the modification would have been to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories and further to accurately detect objects under various scenarios. Since both Rasheed and Zhu share the same endeavor of systems that perform region filtering/normalizing. Wherein Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017], and Zhu’s system improves object detection by accurately detecting objects under various scenarios, see Zhu’s Par. [0002].
Regarding claim 7, Metzner in view of Rasheed further in view of Kao and Jiang teaches the method of claim 3.
However, Metzner in view of Rasheed further in view of Kao and Jiang does not explicitly teach wherein the filtering protocol on the region data includes reducing redundant or overlapping bounding boxes.
Zhu teaches wherein the filtering protocol on the region data includes reducing redundant or overlapping bounding boxes (“or” indicates a selection, therefore, only one of the options is the instant scope of the claim, the examiner selects “reducing redundant bounding boxes”, Par. [0032] discloses “determine whether…excluded from consideration as an object…ignore blobs of insignificant size (e.g., below a threshold number of pixels” indicating a region filtering wherein to remove or ignore boxes/blobs that have insignificant size or being redundant).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed further in view of Kao and Jiang of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Zhu of wherein the filtering protocol on the region data includes reducing redundant or overlapping bounding boxes.
Wherein having Matzner’s method having wherein the filtering protocol on the region data includes reducing redundant or overlapping bounding boxes.
The motivation behind the modification would have been to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories and further to accurately detect objects under various scenarios. Since both Rasheed and Zhu share the same endeavor of systems that perform region filtering/normalizing. Wherein Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017], and Zhu’s system improves object detection by accurately detecting objects under various scenarios, see Zhu’s Par. [0002].
Claims 9-11 are rejected under 35 U.S.C. 103 as being unpatentable over Maximilian Metzner et. al. (“US 2024/0386637 A1” hereinafter as “Metzner”) in view of Hanoona Abdul Rasheed et. al. (“US 2024/0203085 A1” hereinafter as “Rasheed”) further in view of Hao-Cheng Kao et. al. (“US 2020/0065623 A1” hereinafter as “Kao”).
Regarding claim 9, Metzner teaches a system for training (Par. [0002] discloses “method and systems for providing or for generating and providing training image data for training a function”) a multi-task computer vision machine-learning model, comprising: (Pars. [0062-65] discloses “the training data TBD may be used to train a function, for example an object recognition function…training object recognition functions for systems that are used, for example, in the field of autonomous driving” indicating generating of training data for multi functions [analogous to multi-task object recognition functions computer vision machine learning model]; Par. [0032] indicates a multi-task model including labeling, segmenting, annotating, size and position determination): one or more annotation specialist models configured to (Par. [0041] discloses “FIG. 1 depicts a flow chart of a computer-implemented method…depicts possible interim results of the methods steps of the method”, the programmed processors to perform this function is analogous to the recited one or more annotation specialist models as claimed): receive a plurality of images to be annotated (Par. [0051] discloses “a multiplicity of modified images MAB may be generated, that (for example, together with the annotated image AB) are provided as training image data”); and generate pre-filtered annotations for the plurality of images (Par. [0052] discloses “a further image HB. In the image HB, a (partially) unpopulated printed circuit board may be seen, on which none of the objects are depicted”; moreover, Par. [0056] discloses “it may be expedient that the area of the background image HB with which the image are defined by the label L2 is overwritten, and the area to be overwritten itself correspond to one another” therefore, the area in HB image is to be overwritten [analogous to pre-filtered], and the image area of the HB image is defined by a label [analogous to annotation] hence, the image HB with its areas being analogous to pre-filtered annotations for the plurality of images as claimed); a data filtering module configured to (Par. [0041] discloses “FIG. 1 depicts a flow chart of a computer-implemented method…depicts possible interim results of the methods steps of the method”; the programmed processors to perform this function is analogous to the recited a data filtering module as claimed): filter the pre-filtered annotations so as to output candidate annotations for the plurality of images (Par. [0057] discloses “the modified image MAB has been produced, for example, by such overwriting. Such a modified image MAB is highly realistic and improves the quality and accuracy of a function when it is trained” indicating a modified image by overwriting and modifying of the image to produce highly realistic image, hence is a product of filtering since it removes/modifies/overwrite objects from the image [filtering out objects]); an iterative data refinement model configured to (Par. [0041] discloses “FIG. 1 depicts a flow chart of a computer-implemented method…depicts possible interim results of the methods steps of the method”; the programmed processor to perform this step is analogous to the recited iterative data refinement model): iteratively train (Par. [0059] discloses “repeating steps S2 and S3, large training data image sets may be rapidly generated” indicating the generation of the training data to train the trained object recognition model includes repeating steps of annotating/labeling, therefore, the employment of the training of the object recognition model include iteratively annotate each of the received images as claimed) the multi-task computer vision machine-learning model on the plurality of images annotated by the candidate annotations (Pars. [0034-35] discloses “training a function is provided where a (for example, untrained or pretrained) function is trained with training image data…a trained function for use in checking the accuracy of the population of printed circuit boards in the production of printed circuit boards”; Par. [0032] discloses “annotation contains information about the size and position of the at least one object in the annotated image and/or a segmentation, i.e., information about all the pixels associated with the object” indicating a multi-task model including labeling, segmenting, annotating, size and position determination); and a final annotation module (Par. [0041] discloses “FIG. 1 depicts a flow chart of a computer-implemented method…depicts possible interim results of the methods steps of the method”; the programmed processor performing this function is analogous to the recited final annotation module as claimed) configured to store the candidate annotations into the corpus as a final annotation for an associated image of the plurality of images (Pars. [0059-60] discloses “by repeating steps S2 and S3, large training data image sets may be rapidly generated…for n objects…unique variations in the presence of the objects may be generated…different images may therefore be produced from an image 253 for training” indicating generating a corpus [“large training data image sets with unique variations”] of training data with label for these associated images).
However, Metzner does not explicitly teach an enhancement model configured to filter the pre-filtered annotations.
Rasheed teaches an enhancement model configured to filter the pre-filtered annotations (elements of 112f interpretation, Par. [0004] discloses “object detection is a foundation for high-level tasks” indicating the object detection model here is a foundation model being pre-trained; Par. [0017] discloses “object detection using high-quality object proposals from a pre-trained multi-modal vision transformer” indicating a multi-modal vision transformer being the annotator of an enlarged detector vocabulary; Par. [0038] discloses “the present approach connects the image, region, and language representations to generalize better to novel open-vocabulary objects” indicating a large language model, Par. [0090] discloses “L2 normalization is used on the region and text embeddings before computing the RKD loss and final classification scores” indicating a text normalization/filtering according to a region score model; Rasheed’s Par. [0054] discloses “all the boxes are arranged according to their cls scores. Then, a non-maximum suppression is applied with a threshold…all of the bounding boxes…with another bounding boxes are discarded”, wherein discarding indicating blacklisting).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, a data module configured to filter the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Rasheed of having the data module being a data filtering and enhancement module.
Wherein Matzner’s method having wherein the data module being a data filtering and enhancement module.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories. Since both Metzner and Rasheed share the same endeavor of systems that perform object detection and labelling in images. Wherein Metzner’s system improve generating training data with labels more effectively and less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], and Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017].
However, Metzner in view of Rasheed does not explicitly teach filtering the pre-filtered annotations in accordance with predefined noise criteria; store a candidate annotation of the candidate annotations.
Kao teaches filtering the pre-filtered annotations in accordance with predefined noise criteria (Par. [0054-55] discloses “determining the inter-annotator consistency…if all of the consistencies are higher than a threshold, the labelled results may be determined to be valid for training the AI machine and be fed to the AI machine”, moreover, this consistency filtering is the filtering step in FIG. 10B illustrates, moreover, Par. [0051] discloses “since the annotator consistently label the three raw data as the class C, the processor may obtain a high intra-annotator consistency of the annotator after calculation the intra-class correlation coefficient”; therefore, indicating a filtering step of the labels [annotations] according to a threshold [predetermined noise criteria as claimed]); store a candidate annotation of the candidate annotations (Par. [0115] discloses “after the annotators finish their labeling operations, the first raw data with the shown bounding regions (i.e., labelled data) may be referred to as a first labelled result and retrieved by the processor. With the first labelled result, the processor may accordingly determine whether the first labelled result is valid for training…based on a plurality of consistencies of the annotators” indicating that for each of the label produced by the annotators, only when the consistencies between the annotators in their labelling meet certain criteria then, that label is used for training data [analogous to storing the candidate annotation into the corpus as a final annotation for its associated image]).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed of having a system for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image, or (2) adding the candidate annotation to its associated image using the one or more annotation specialist models and the data filtering and enhancement module for subsequent iterative annotation and filtering, with the teachings of Kao of filtering the pre-filtered annotations in accordance with predefined noise criteria; store a candidate annotation of the candidate annotations.
Wherein Matzner’s system wherein having filtering the pre-filtered annotations in accordance with predefined noise criteria; store a candidate annotation of the candidate annotations.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to verify generated labelled data to obtain more accurate and robust training data. Since both Metzner and Kao share the same endeavor of systems that perform training data generation. Wherein Metzner’s system improve generating training data with labels more effectively and less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], and Kao’s system improves generating of training data by verifying generated labelled data to obtain more accurate and robust training data, see Kao’s Pars. [0003-0004].
Regarding claim 10, Metzner in view of Rasheed and Kao teaches the system of claim 9.
However, Metzner does not explicitly teach where the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model.
Kao teaches where the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model (“one or more” indicates a selection, therefore, only one of these options is the instant scope of the claim, the examiner selects “trained object proposal and detection models”, wherein Kao’s Par. [0042] discloses “the annotator may recognize the raw data as an image with a cat” indicating an object determination model, moreover, Par. [0060] discloses “certain annotator back to be trained again” indicating the annotators are trained models).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed and Kao of having a system for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image, or (2) adding the candidate annotation to its associated image using the one or more annotation specialist models and the data filtering and enhancement module for subsequent iterative annotation and filtering, with the teachings of Kao of where the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model.
Wherein Matzner’s system having where the one or more annotation specialist models are trained models including one or more of a (1) trained caption model; (2) trained grounding model; (3) trained segmentation model; (4) trained object proposal and detection models; and (5) trained optical character recognition model.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to verify generated labelled data to obtain more accurate and robust training data. Since both Metzner and Kao share the same endeavor of systems that perform training data generation. Wherein Metzner’s system improve generating training data with labels more effectively and less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], and Kao’s system improves generating of training data by verifying generated labelled data to obtain more accurate and robust training data, see Kao’s Pars. [0003-0004].
Regarding claim 11, Metzner in view of Rasheed and Kao teaches the system of claim 9.
However, Metzner in view of Kao does not explicitly teach where the filtering and enhancement model comprises one or more of a (1) text filter, (2) enhancement model, and (3) region filtering model
Rasheed teaches where the filtering and enhancement model comprises one or more of a (1) text filter, (2) enhancement model, and (3) region filtering model (“one or more of” indicates selection, therefore, only one of the options is the instant scope of the claim, the examiner selects “region filtering model”, wherein Rasheed’s Par. [0090] discloses “L2 normalization is used on the region and text embeddings” indicating a text normalization or filtering; moreover, Par. [0086] discloses “a dataset for large vocabulary instance segmentation…for pseudo-labeling process” indicates a model for object detection based on segmentation and object pseudo-labeling which is analogous to the model of Metzner, hence it’s obvious to modify Metzner’s object detection model to further include region and text embedding segmentation and normalization for labeling task of images).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed and Kao of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Rasheed of where the filtering and enhancement model comprises one or more of a (1) text filter, (2) enhancement model, and (3) region filtering model.
Wherein Matzner’s method having where the filtering and enhancement model comprises one or more of a (1) text filter, (2) enhancement model, and (3) region filtering model.
The motivation behind the modification would have been to generate training data with labels more effectively and less errors by using real and non-synthetic training image and further to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories. Since both Metzner and Rasheed share the same endeavor of systems that perform object detection and labelling in images. Wherein Metzner’s system improve generating training data with labels more effectively and less errors by using real and non-synthetic training image, see Metzner’s Par. [0009] and Par. [0013], and Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017].
Claims 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Maximilian Metzner et. al. (“US 2024/0386637 A1” hereinafter as “Metzner”) in view of Hanoona Abdul Rasheed et. al. (“US 2024/0203085 A1” hereinafter as “Rasheed”) and Hao-Cheng Kao et. al. (“US 2020/0065623 A1” hereinafter as “Kao”) and Zhaowen Wang et. al. (“US 2017/0200065 A1” hereinafter as “Wang”).
Regarding claim 12, Metzner in view of Rasheed and Kao teaches the system of claim 9.
However, Metzner in view of Rasheed and Kao does not explicitly teach wherein the final annotation for the associated image include at least a brief caption, a detailed caption, and a more detailed caption.
Wang teaches wherein the final annotation for the associated image include at least a brief caption, a detailed caption, and a more detailed caption (Par. [0046] discloses “the weak annotations provide detail information regarding images at a deeper level of understanding…the weak annotations are relied upon to generate a collection of keywords…of low-level image details” therefore, the annotation of image include low-level detail [brief caption] that can be used to provide detailed information about the image [detailed action] of deeper level [a more detailed caption]).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Metzner in view of Rasheed and Kao of having a method for annotating images to create a corpus for training a multi-task computer vision machine learning model, comprising: receiving, a plurality of images to be annotated, generating pre-filtered annotations for the plurality of images, via a data filtering and enhancement module filtering the pre-filtered annotations so as to output candidate annotations for the plurality of images, with the teachings of Wang of wherein the final annotation for the associated image include at least a brief caption, a detailed caption, and a more detailed caption.
Wherein Matzner’s method having wherein the final annotation for the associated image include at least a brief caption, a detailed caption, and a more detailed caption.
The motivation behind the modification would have been to perform object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories and further to generate precise and complex image captions of content of image based on associated captions. Since both Rasheed and Zhu share the same endeavor of systems that perform region filtering/normalizing. Wherein Rasheed’s system improves object detection with high-quality proposals in order to enlarge detector vocabulary and generalize toward novel object categories, see Rasheed’s Par. [0017], and Wang’s system improves generating precise and complex image captions of content of image based on associated captions, see Wang’s abstract.
Regarding claim 13, Metzner in view of Rasheed and Kao teaches the system of claim 12, Metzner teaches wherein the final annotation for the associated image is associated with one or more of a detected object and a region of the associated image (Metzner’s Par. [0059] discloses “by repeating steps S2 and S3, large training data image sets may be rapidly generated” indicating generating a corpus of training data with label for these associated images, moreover, Par. [0052] discloses “further image…on which none of the objects (to be detected)…are depicted” indicating a plurality of objects to be detected, Abstract discloses “the annotation describing an image region in which the at least one object is contained”).
Regarding claim 14, Metzner in view of Rasheed and Kao teaches the system of claim 13, Metzner teaches wherein the final annotation for the associated image is one of a plurality of final annotations for the associated image that includes at least a (1) text annotation, (2) region-text pair annotation, and (3) text-phrase-region triplet annotation (“is one of a plurality of” indicates selection, therefore, only one of the options is the instant scope of the claim, the examiner selects “text annotation”, wherein Metzner’s Abstract discloses “the annotation describing an image region in which the at least one object is contained”).
Pertinent Prior Art(s)
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Seybold, Bryan Andrew et. al., “US 2025/0209794 A1”, teaches techniques for improving the performance of video retrieval systems and audio retrieval systems are described herein. A computing system can obtain a captioned image with an associated caption and a first video having a plurality of frames. Additionally, the system can determine a feature vector of the captioned image and a feature vector of a first frame in the plurality of frames. Moreover, the system can calculate a similarity value between the captioned image and the first frame based on the feature vector of the captioned image and the feature vector of the first frame. Furthermore, the system can transfer the associated caption to the first frame based on the similarity value. Subsequently, the system can generate a video clip based on the first frame. The system can also store and index the video clip in a video captioning database.
Han, Seung Ho et. al., “US 12437565 B2”, teaches an apparatus and method for automatically generating an image caption is provided capable of giving an explanation by using Bayesian inference and an image area-word mapping module on the basis of a deep learning algorithm. An apparatus for automatically generating an image caption, according to one embodiment of the present invention, includes: an automatic caption generation module for creating a caption by applying a deep learning algorithm to an image received from a client; a caption basis generation module for creating a basis for the caption by mapping a partial area in the image received from the client with respect to important words in the caption received from the automatic caption generation module; and a visualization module for visualizing the caption received from the automatic caption generation module and the basis for the caption received from the caption basis generation module to return same to the client.
Stubler, Peter O. et. al., “US 2002/0188602 A1”, teaches a method of generating captions or semantic labels for an acquired image is based upon similarity between the acquired image and one or more stored images that are maintained in an image database environment, where the stored images have preexisting captions or labels associated with them. The method includes the steps of: a) acquiring an image for evaluation with respect to the stored images; b) automatically extracting metadata from the acquired image without requiring user interaction in the image database environment; c) automatically selecting one or more stored images having metadata similar to the extracted metadata, thereby providing selected stored images with metadata similar to the acquired image; and d) generating captions or labels for the acquired image from the preexisting captions or labels associated with the selected stored images.
Bodie, Jeffrey C., “US 2006/0092291 A1”, teaches a digital imaging system includes facilities to capture an image and a related audio annotation, convert the audio annotation to text by voice recognition, and associate and edit the related image, audio annotation, and image caption.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHUONG HAU CAI whose telephone number is (571)272-9424. The examiner can normally be reached M-F 8:30 am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHUONG HAU CAI/ Examiner, Art Unit 2673
/CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673