Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Use of indicates a limitation is not explicitly disclosed by the reference alone.
Claim(s) 1-4, 8, 10-13, 17, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang, HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes in view of Han, A Survey on Vision Transformer.
Claim 1
Wang discloses a method, executed by at least one processor, comprising:
receiving an input comprising:
a 3D point cloud of a scene comprising a goal object, and a text comprising a natural language instruction associated with the goal object (Wang, Page 5: “The condition module takes input from two modalities (i.e., the given scene S and the language description L)”);
applying a text tokenizer to the text to obtain a tokenized text (Wang, Page 5: “The language description is a tokenized word sequence of length D, denoted as L1:D = [w1,··· ,wD]”);
generating first scene features by application of a pre-trained (Wang, Page 7: “For each scene, we down-sample N = 32768 points from the original scanned scene and obtain N′ = 128 points with a 512-D feature for each point after applying the point transformer”);
down sampling the first scene features to obtain second scene features (Wang, Page 7: “For each scene, we down-sample N = 32768 points from the original scanned scene and obtain N′ = 128 points with a 512-D feature for each point after applying the point transformer”);
obtaining a conditional latent based on a fusion of the second scene features with the text features (Wang, Page 5; “The scene and language features are finally concatenated and mapped to a conditional latent embedding zc with FC layers”);
predicting a sequence of motion parameters for a motion of a parametric human body model towards the goal object for a specific time duration by applying a conditional motion generator on the conditional latent (Wang, Page 6: “Motion encoder We first use a bidirectional GRU to obtain a sequence-level feature of the input motion Θ1:T. Next, the output is concatenated with the conditional embedding zc, followed by an MLP layer to predict the Gaussian distribution parameters (i.e., µ and Σ). Finally, we sample a latent vector z”); and
obtaining 3D human meshes for a plurality of motion frames based on the sequence of motion parameters and the parametric human body model (Wang, Page 6: “Motion decoder Following Petrovich et al. [2021], we use a transformer decoder to generate a sequence of parameters Θ1:T for a given duration T. Specifically, we use T sinusoidal positional embeddings to query the concatenation of the sampled latent z and the conditional embedding zc. The outputs of the transformer decoder are mapped into body meshes with the differentiable SMPL-X model.”).
Wang does not explicitly disclose, but Han discloses generating text features by applying a text encoder of a pre-trained vision- language model on the tokenized text (e.g. CLIP; Han, Section 3.5: “CLIP learns both text and image embeddings jointly to maximize the cosine similarity of those N matched embeddings while minimize N2−N incorrectly matched embeddings”);
Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to use a pre-trained model like CLIP.
One of ordinary skill in the art would have motivation because “Contrastive Language-Image Pre-training (CLIP) [40] takes natural language as supervision to learn more efficient image representation.”(Han, Section 3.5). One of ordinary skill in the art would have had a reasonable expectation of success because both references consider tokenized text to supervise image output.
Wang does not explicitly disclose, but Han discloses U-net (“There are also other types of architectures, such as two-stream architecture [79] and U-net architecture”);
Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to use U-net
One of ordinary skill in the art would have motivation for improved architecture. One of ordinary skill in the art would have had a reasonable expectation of success because both references consider tokenized text to supervise image output.
Claim 2
Wang does not disclose, but Han discloses wherein the pre-trained vision-language model is a Contrastive Language-Image Pre-Training (CLIP) model (e.g. CLIP; Han, Section 3.5: “CLIP learns both text and image embeddings jointly to maximize the cosine similarity of those N matched embeddings while minimize N2−N incorrectly matched embeddings”);
Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to use a pre-trained model like CLIP.
One of ordinary skill in the art would have motivation because “Contrastive Language-Image Pre-training (CLIP) [40] takes natural language as supervision to learn more efficient image representation.”(Han, Section 3.5). One of ordinary skill in the art would have had a reasonable expectation of success because both references consider tokenized text to supervise image output.
Claim 3
Wang discloses wherein the pre-trained U-Net scene encoder is a Point Transformer-based neural network (Wang, Page 7: “For each scene, we down-sample N = 32768 points from the original scanned scene and obtain N′ = 128 points with a 512-D feature for each point after applying the point transformer”);
Claim 4
Wang discloses further comprising feeding position and color information of each 3D point of the 3D point cloud to the pre-trained U-Net scene encoder to generate the first scene features which include a point feature vector for each 3D point of the 3D point cloud (3D point cloud encodes position; Color corresponds to value at each point; “We denote the given scene as S ∈ RN×6, representing an RGB-colored point cloud of N points…”).
Claim 8
Wang discloses wherein the fusion of the second scene features with the text features comprises:
concatenating the second scene features with the text features to obtain a concatenated feature (Wang, Section 4.4: “map the obtained 768-D word-level features into 512-D before concatenating them with the point features…”); and
applying a self-attention layer on the concatenated feature to obtain a fused feature, wherein the conditional latent is generated based on the fused features (Wang, Section 4.4: “Instead of feeding the scene and the language feature into the self-attention layer, we directly concatenate them as the global conditional feature. We denote this model as w/o self-att. We train and test on the walk subset.”)
Claim 10
Examiner’s Interpretation:
Machine readable media can encompass forms of signal transmission media that falls outside of the four statutory categories of invention. MPEP 2106; citing In re Nuijten, 500 F.3d 1346, 84 USPQ2d 1495 (Fed. Cir. 2007). A claim whose BRI covers both statutory and non-statutory embodiments embraces subject matter that is not eligible for patent protection and therefore is directed to non-statutory subject matter. MPEP 2106.
Claims 10-18 as drafted recite One or more non-transitory computer-readable storage media.
The broadest reasonable interpretation of the claimed medium in view of Applicant’s specification covers only eligible subject matter.
Claim Mapping:
The same teachings and rationales in claim 1 are appliable to claim 10.
Claim 11
The same teachings and rationales in claim 2 are appliable to claim 11.
Claim 12
The same teachings and rationales in claim 3 are appliable to claim 12.
Claim 13
The same teachings and rationales in claim 4 are appliable to claim 13.
Claim 17
The same teachings and rationales in claim 8 are appliable to claim 17.
Claim 19
The same teachings and rationales in claim 1 are appliable to claim 19, with a CPU/GPU based system disclosing a corresponding system as claimed.
Claim(s) 7, 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang, HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes in view of Han, A Survey on Vision Transformer and Choe, PointMixer: MLP-Mixer for Point Cloud Understanding
Claim 7
Wang does not explicitly disclose but Choe discloses wherein the down sampling is performed using a k-nearest neighbor classifier (Choe, Section 2: “this paper adopts k-Nearest Neighbor (kNN) for local neighborhood sampling…While downsampling layers adopt pooling with kNN and FPS, upsampling layers re-compute kNN”)
Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to downsample with KNN.
Wang already utilizes downsampling of the point representation, and one of ordinary skill in the art would have motivation to apply the operation to point samples (“or various 3D perception tasks such as object shape classification [80], semantic segmentation [2] and point cloud reconstruction tasks”)(Choe, Section2). One of ordinary skill in the art would have had a reasonable expectation of success because both references consider application to point sampling.
Claim 16
The same teachings and rationales in claim 7 are appliable to claim 16.
Allowable Subject Matter
Claim(s) 5, 6, 9, 14, 15, 18, 20 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding claim(s) 5, 14, 20 Han, cited to disclose U-net does not suggest application of U-net to “obtaining the pre-trained U-Net scene encoder by pre-training the U-Net scene encoder until a distance between the image feature vectors and the point feature vector is a minimum”
Regarding claim(s) 6, 15, Choe, cited to disclose Knn, does not suggest applying an average pooling operation on the set of k-nearest neighboring vectors around each point feature vector of the set of point feature vectors to obtain a plurality of average pooled vectors,wherein the second scene features include the plurality of average pooled vectors
Regarding claim(s) 9, 18 Wang considers regularlized losses but uses position rather than size and does not disclose two regularization losses associated with a category of the goal object and a size of the goal object.
Additional Prior Art
Additional prior art relevant to Applicant’s disclosure but not relied upon:
Yi (US 2025/0232506) also discloses a goal state:
PNG
media_image1.png
487
610
media_image1.png
Greyscale
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN M GRAY whose telephone number is (571)272-4582. The examiner can normally be reached on Monday through Friday, 9:00am-5:30pm (EST).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached on (571)272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN M GRAY/Primary Examiner, Art Unit 2611