Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 09/26/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim 14 is objected to because of the following informalities:
Claim 14 recites “one or more processors configure to:” in line 2. The examiner believes it should recite “one or more processors configured to:”. Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
Position embedder in claims 5 and 18.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature “Masked Autoencoders for Point Cloud Self-supervised Learning” by Pang et al. (hereinafter Pang), dated 03/28/2022, available at https://doi.org/10.48550/arXiv.2203.06604, in view of us patent application publication no. 2024/0394546 to Cho et al. (hereinafter Cho).
For claim 10, Pang as applied discloses a processor-implemented method comprising:
masking at least a part of voxel data obtained from a point cloud to generate masked voxels (see, e.g., par. 6 in sec. 1, par. 1 in sec. 3, pars. 1, 2, 3 in sec. 3.1, which teach generating point patches by dividing input point cloud and masking the point patches at a given ratio);
obtaining feature information about the masked voxels from unmasked voxels, which are not masked, through a backbone network that is pre-trained (see, e.g., par. 6 in sec. 1, par. 1 in sec. 3, pars. 4-5 in sec. 3.1, pars. 1-4 in sec. 3.2, and par. 2 in sec. 4.1 and FIGS. 1, 2 and 4, which teach embedding visible/unmasked patches into visible tokens and obtaining decoded mask tokens from the visible tokens; the examiner notes that while the teaching is described with respect to a pre-training but it would also be applicable to a pre-trained model because a trained model behaves as it was trained);
wherein the backbone network that is pre-trained is a network in which a parameter is trained based on a similarity between a masking feature vector for the masked voxels extracted through a teacher network and a prediction feature vector extracted through the backbone network (see, e.g., par. 6 in sec. 1, par. 5 in sec. 3.2 and par. 1 in sec. 3.3, which teach that the autoencoder’s backbone is trained based on a l2 similarity/reconstruction loss between the predicted masked point patches from the prediction head of the backbone and the ground truth masked point patches).
Pang as applied does not explicitly teach calculating a loss between the predicted feature vectors and the vectors from a teach network. Cho in the analogous art teaches training the backbone network by calculating a difference between the feature vectors of teacher and student networks and optimizing parameters of the student network (see, e.g., pars. 29, 39-40, 50-55, and 61-63 and FIGS. 1, 3 and 4 of Cho).
It would have been obvious to one of ordinary skill in the art to modify Pang to calculate the similarity using the feature vectors from a teacher network as taught by Cho because doing so would enhance the effect of learning by knowledge distillation between the backbone and the teacher (see pars. 6, 19 and 63 of Cho).
Claim(s) 11 and 12 is is/are rejected under 35 U.S.C. 103 as being unpatentable over Pang in view of Cho and further in view of non-patent literature titled “Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling” by Yu et al. (hereinafter Yu), dated 06/06/2022, available at https://doi.org/10.48550/arXiv.2111.14819.
For claim 11, while Pang in view of Cho and Deva does not teach, Yu in the analogous art teaches that the feature information comprises a semantic feature and a geometric feature of the masked voxels (see, e.g., par. 5 in sec. 3.3, which teaches that that the predicted point tokens include information about geometric structure and semantic patterns).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho to include geometric and semantic features of the masked point patches as taught by Yu because doing so would allow the model to understand the high-level semantics of point clouds (see, e.g., par. 5 in sec. 3.3 of Yu).
For claim 12, while Pang in view of Cho does not teach, Yu in the analogous art teaches that the parameter of the backbone network that is pre-trained is trained based on geometric information about the masked voxels and a loss function between the geometric information about the masked voxels and geometric information obtained from the point cloud (see, e.g., pars. 4 and 5 in sec. 3.3, which teach training the backbone transformers using a contrastive loss function between the features of the point tokens corresponding to the masked locations and the features of the tokens corresponding to the virtual samples in the point cloud, wherein the point patches/tokens include geometric structure information).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho to include geometric and semantic features of the masked point patches as taught by Yu because doing so would allow the model to understand the high-level semantics of point clouds (see, e.g., par. 5 in sec. 3.3 of Yu).
Claim(s) 1, 4, 5, 8-9, 13, 14, and 17-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Pang in view of Cho and further in view of us patent application publication no. 2025/0061699 to Devaguptapu et al. (hereinafter Deva).
For claims 1, 9 and 14, Pang as applied teaches a processor-implement method (see, e.g., abstract) comprising:
masking at least a part of voxel data obtained from a point cloud to generate masked voxels (see, e.g., par. 6 in sec. 1, par. 1 in sec. 3, pars. 1-3 in sec. 3.1 and FIGS. 1 and 3, which teach generating point patches by dividing the input point cloud and masking the point patches at a given ratio);
obtaining feature information about the masked voxels from unmasked voxels, which are not masked, through a backbone network (see, e.g., par. 6 in sec. 1, par. 1 in sec. 3, pars. 4-5 in sec. 3.1, pars. 1-4 in sec. 3.2 and FIGS. 1 and 3, which teach embedding visible/unmasked patches into visible tokens and obtaining decoded mask tokens from the visible tokens);
extracting a prediction feature vector from the feature information through a feature prediction model (see, e.g., par. 5 in sec. 3.2 and par. 1 in sec. 3.3, which teach obtaining predicted masked point patches from the prediction head);
extracting a masking feature vector for the masked voxels (see, e.g., par. 2 in sec. 3.1 and par. 1 in sec. 3.3, which teach obtaining the ground truth masked point patches); and
training the backbone network by updating the parameter of the backbone network based on a similarity between the prediction feature vector and the masking feature vector (see, e.g., par. 6 in sec. 1, par. 5 in sec. 3.2 and par. 1 in sec. 3.3 and FIGS. 1 and 3, which teach that the autoencoder’s backbone is trained based on a l2 similarity/reconstruction loss between the predicted masked point patches from the prediction head of the backbone and the ground truth masked point patches).
Pang as applied does not explicitly teach that the ground truth masked point patches are extracted through the teacher module and that the calculated loss is between the predicted feature vectors and the feature vectors extracted from the teacher module.
Cho in the analogous art teaches extracting a feature vector through a teacher network (see, e.g., pars. 29, 32 and 63 and FIG. 1 of Cho) and training the backbone network by calculating a difference between the feature vectors of teacher and student networks and optimizing parameters of the student network (see, e.g., pars. 29, 39-40, 50-55, and 61-63 and FIGS. 1, 3 and 4 of Cho).
It would have been obvious to one of ordinary skill in the art to modify Pang to calculate the similarity using the feature vectors from a teacher network as taught by Cho because doing so would enhance the effect of learning by knowledge distillation between the backbone and the teacher (see pars. 6, 19 and 63 of Cho).
Pang in view of Cho does not explicitly teach “determining a parameter of a teacher module based on a parameter of the backbone network.” Deva in the analogous art teaches updating parameters of the teacher network after a certain number of updates to the student network (see, e.g., pars. 19 and 29 of Deva; in view of pars. 46-47 in the specification, which describe the relationship between the backbone network and the teacher network, the examiner interprets the student network correspond to the backbone network).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho to update the teacher network’s parameters as taught by Deva because doing so would yield predictable results of allowing an efficient tracking of student parameters (see MPEP 2143(I)(D)).
For claims 4 and 17, Pang in view of Cho does not explicitly teach that the determining of the parameter of the teacher module based on the parameter of the backbone network comprises updating the parameter of the teacher module at a predetermined interval using the parameter of the backbone network.
Deva in the analogous art teaches updating parameters of the teacher network after a certain number of updates to the student network (see, e.g., pars. 19 and 29 of Deva; in view of pars. 46-47 in the specification, which describe the relationship between the backbone network and the teacher network, the examiner interprets the student network to correspond to and read on the backbone network).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho to update the teacher network’s parameters as taught by Deva because doing so would yield predictable results of allowing an efficient tracking of student parameters (see MPEP 2143(I)(D)).
For claims 5 and 18, Pang in view of Cho and Deva teaches that the feature prediction model comprises:
a position embedder for extending a dimension of the feature information about the masked voxels (see, e.g., par. 3 in sec. 3.1 of Pang, which teaches extending a dimension of the masked point patch with a mask token); and
a transformer-based prediction model for extracting the prediction feature vector for the feature information of the extended dimension (see, e.g., par. 2 in sec. 3.2 of Pang, which teaches obtaining the decoded mask tokens by adding the positional embeddings to the transformer blocks).
For claim 8, Pang in view of Cho and Deva teaches masking at least a part of other voxel data obtained from another point cloud (see, e.g., par. 6 in sec. 1, par. 1 in sec. 3, pars. 1-3 in sec. 3.1 of Pang, which teach generating point patches by dividing input point cloud and masking the point patches at a given ratio); and obtaining other feature information about other masked voxels from other unmasked voxels, which are not masked, through the trained backbone network (see, e.g., par. 6 in sec. 1, par. 1 in sec. 3, pars. 4-5 in sec. 3.1, pars. 1-4 in sec. 3.2, and par. 2 in sec. 4.1, and pars. 1-4 in sec. 4.2 and FIGS. 1, 2 and 4 of Pang, which teach embedding visible/unmasked patches into visible tokens, obtaining decoded mask tokens from the visible tokens, and using the trained backbone network with different point clouds for downstream tasks).
For claim 13, Pang in view of Cho does not explicitly teach that a parameter of the teacher network is updated at a predetermined interval through an exponential moving average (EMA) of the parameter of the backbone network.
Deva in the analogous art teaches updating parameters of the teacher network after a certain number of updates to the student network (see, e.g., pars. 19 and 29 of Deva; in view of pars. 46-47 in the specification, which describe the relationship between the backbone network and the teacher network, the examiner interprets the student network correspond to the backbone network).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho to update the teacher network’s parameters as taught by Deva because doing so would yield predictable results of allowing an efficient tracking of student parameters (see MPEP 2143(I)(D)).
Claim(s) 2, 3, 7, 15, 16, and 20 is is/are rejected under 35 U.S.C. 103 as being unpatentable over Pang in view of Cho and Deva and further in view of Yu.
For claims 2 and 15, while Pang in view of Cho and Deva does not teach, Yu in the analogous art teaches that the obtaining of the feature information about the masked voxels from the unmasked voxels, which are not masked, through the backbone network comprises obtaining geometric information from the unmasked voxels from the backbone network (see, e.g., par. 5 in sec. 3.3, which teaches that that the predicted point tokens include information about geometric structure and semantic patterns).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho and Deva to include geometric and semantic features of the masked point patches as taught by Yu because doing so would allow the model to understand the high-level semantics of point clouds (see, e.g., par. 5 in sec. 3.3 of Yu).
For claims 3 and 16, while Pang in view of Cho and Deva does not explicitly teach, Yu in the analogous art teaches: obtaining the geometric information about the masked voxels through a geometric prediction model (see, e.g., par. 5 in sec. 3.3., which teach extracting features through the momentum feature encoder); and training the parameter of the backbone network based on a loss function between the geometric information about the masked voxels and geometric information obtained from the point cloud (see, e.g., pars. 4 and 5 in sec. 3.3 of Yu, which teach training the backbone transformers using a contrastive loss function between the features of the point tokens corresponding to the masked locations and the features of the tokens corresponding to the virtual samples in the point cloud, wherein the point patches/tokens include geometric structure information).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho and Deva to include geometric and semantic features of the masked point patches as taught by Yu because doing so would allow the model to understand the high-level semantics of point clouds (see, e.g., par. 5 in sec. 3.3 of Yu).
For claims 7 and 20, while Pang in view of Cho and Deva does not explicitly teach, Yu in the analogous art teaches that the feature information comprises a semantic feature of the masked voxels (see, e.g., par. 5 in sec. 3.3, which teaches that that the point token corresponding to the masked locations includes information about geometric structure and semantic patterns).
It would have been obvious to one of ordinary skill in the art to modify Pang in view of Cho to include geometric and semantic features of the masked point patches as taught by Yu because doing so would allow the model to understand the high-level semantics of point clouds (see, e.g., par. 5 in sec. 3.3 of Yu).
Allowable Subject Matter
Claims 6 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
In regard to claim 6 and 19, when considered as a whole, prior art of record fails to disclose or render obvious, alone or in combination:
“the extracting of the prediction feature vector from the feature information through the feature prediction model comprises:
extending a dimension for each position for the feature information about the unmasked voxels and a token about the masked voxels; and
extracting the prediction feature vector by entering the position, in which the dimension is extended, into the feature prediction model.”
Additional Citations
The following table lists several references that are relevant to the subject matter claimed and disclosed in this Application. The references are not relied on by the Examiner, but are provided to assist the Applicant in responding to this Office action.
Citation
Relevance
Liao et al. (us pat. pub. 2024/0233134)
Describes 3D medical image segmentation under weak supervision. Particularly, the disclosure relates to a weakly-supervised segmentation framework for segmenting 3D medical images, where the framework utilizes information on geometric prior and contrastive similarity in 3D medical image segmentation.
Goncharuk et al. (us pat. pub. 2024/0070459)
Describes a system and method for effectively training a machine learning model to identify features in DAS and/or seismic imaging data with limited or no human labels. This is accomplished using a masked autoencoder (MAE) network that is trained in multiple stages. The first stage is a self-supervised learning (SSL) stage where the model is generically trained to predict data that has been removed (masked) from an original dataset. The second stage involves performing additional predictive training on a second dataset that is specific to a particular geographic region, or specific to a certain set of desired features. The model is fine-tuned using labeled data in order to develop feature extraction capabilities.
Zhang et al. (us pat. pub. 2024/0037852)
Describes a method for reconstructing three-dimensional faces. The method includes: estimating a dynamic reconstruction parameter of a current video frame for three-dimensional face reconstruction by inputting, in response to a steady-state reconstruction parameter of the current video frame for the three-dimensional face reconstruction having been estimated by a pre-constructed teacher network model, the current video frame into a student network model distilled from the teacher network model; and reconstructing a three-dimensional face corresponding to the current video frame by inputting the steady-state reconstruction parameter and the dynamic reconstruction parameter into a pre-constructed three-dimensional deformation model.
Table 1
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See table 1 and form 892.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WOO RHIM whose telephone number is (571)272-6560. The examiner can normally be reached Mon - Fri 9:30 am - 6:00 pm et.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at 571-272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WOO C RHIM/Examiner, Art Unit 2676