Prosecution Insights
Last updated: September 17, 2026
Application No. 19/049,508

GROUNDED HUMAN MOTION GENERATION WITH OPEN VOCABULARY SCENE-AND-TEXT CONTEXTS

Non-Final OA §103
Filed
Feb 10, 2025
Priority
Mar 28, 2024 — provisional 63/571,353
Examiner
GRAY, RYAN M
Art Unit
Tech Center
Assignee
Fuijtsu Limited
OA Round
1 (Non-Final)
88%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
599 granted / 684 resolved
+27.6% vs TC avg
Moderate +12% lift
Without
With
+11.8%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
25 currently pending
Career history
706
Total Applications
across all art units

Statute-Specific Performance

§101
7.6%
-32.4% vs TC avg
§103
70.9%
+30.9% vs TC avg
§102
7.2%
-32.8% vs TC avg
§112
4.1%
-35.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 684 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Use of indicates a limitation is not explicitly disclosed by the reference alone. Claim(s) 1-4, 8, 10-13, 17, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang, HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes in view of Han, A Survey on Vision Transformer. Claim 1 Wang discloses a method, executed by at least one processor, comprising: receiving an input comprising: a 3D point cloud of a scene comprising a goal object, and a text comprising a natural language instruction associated with the goal object (Wang, Page 5: “The condition module takes input from two modalities (i.e., the given scene S and the language description L)”); applying a text tokenizer to the text to obtain a tokenized text (Wang, Page 5: “The language description is a tokenized word sequence of length D, denoted as L1:D = [w1,··· ,wD]”); generating first scene features by application of a pre-trained (Wang, Page 7: “For each scene, we down-sample N = 32768 points from the original scanned scene and obtain N′ = 128 points with a 512-D feature for each point after applying the point transformer”); down sampling the first scene features to obtain second scene features (Wang, Page 7: “For each scene, we down-sample N = 32768 points from the original scanned scene and obtain N′ = 128 points with a 512-D feature for each point after applying the point transformer”); obtaining a conditional latent based on a fusion of the second scene features with the text features (Wang, Page 5; “The scene and language features are finally concatenated and mapped to a conditional latent embedding zc with FC layers”); predicting a sequence of motion parameters for a motion of a parametric human body model towards the goal object for a specific time duration by applying a conditional motion generator on the conditional latent (Wang, Page 6: “Motion encoder We first use a bidirectional GRU to obtain a sequence-level feature of the input motion Θ1:T. Next, the output is concatenated with the conditional embedding zc, followed by an MLP layer to predict the Gaussian distribution parameters (i.e., µ and Σ). Finally, we sample a latent vector z”); and obtaining 3D human meshes for a plurality of motion frames based on the sequence of motion parameters and the parametric human body model (Wang, Page 6: “Motion decoder Following Petrovich et al. [2021], we use a transformer decoder to generate a sequence of parameters Θ1:T for a given duration T. Specifically, we use T sinusoidal positional embeddings to query the concatenation of the sampled latent z and the conditional embedding zc. The outputs of the transformer decoder are mapped into body meshes with the differentiable SMPL-X model.”). Wang does not explicitly disclose, but Han discloses generating text features by applying a text encoder of a pre-trained vision- language model on the tokenized text (e.g. CLIP; Han, Section 3.5: “CLIP learns both text and image embeddings jointly to maximize the cosine similarity of those N matched embeddings while minimize N2−N incorrectly matched embeddings”); Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to use a pre-trained model like CLIP. One of ordinary skill in the art would have motivation because “Contrastive Language-Image Pre-training (CLIP) [40] takes natural language as supervision to learn more efficient image representation.”(Han, Section 3.5). One of ordinary skill in the art would have had a reasonable expectation of success because both references consider tokenized text to supervise image output. Wang does not explicitly disclose, but Han discloses U-net (“There are also other types of architectures, such as two-stream architecture [79] and U-net architecture”); Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to use U-net One of ordinary skill in the art would have motivation for improved architecture. One of ordinary skill in the art would have had a reasonable expectation of success because both references consider tokenized text to supervise image output. Claim 2 Wang does not disclose, but Han discloses wherein the pre-trained vision-language model is a Contrastive Language-Image Pre-Training (CLIP) model (e.g. CLIP; Han, Section 3.5: “CLIP learns both text and image embeddings jointly to maximize the cosine similarity of those N matched embeddings while minimize N2−N incorrectly matched embeddings”); Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to use a pre-trained model like CLIP. One of ordinary skill in the art would have motivation because “Contrastive Language-Image Pre-training (CLIP) [40] takes natural language as supervision to learn more efficient image representation.”(Han, Section 3.5). One of ordinary skill in the art would have had a reasonable expectation of success because both references consider tokenized text to supervise image output. Claim 3 Wang discloses wherein the pre-trained U-Net scene encoder is a Point Transformer-based neural network (Wang, Page 7: “For each scene, we down-sample N = 32768 points from the original scanned scene and obtain N′ = 128 points with a 512-D feature for each point after applying the point transformer”); Claim 4 Wang discloses further comprising feeding position and color information of each 3D point of the 3D point cloud to the pre-trained U-Net scene encoder to generate the first scene features which include a point feature vector for each 3D point of the 3D point cloud (3D point cloud encodes position; Color corresponds to value at each point; “We denote the given scene as S ∈ RN×6, representing an RGB-colored point cloud of N points…”). Claim 8 Wang discloses wherein the fusion of the second scene features with the text features comprises: concatenating the second scene features with the text features to obtain a concatenated feature (Wang, Section 4.4: “map the obtained 768-D word-level features into 512-D before concatenating them with the point features…”); and applying a self-attention layer on the concatenated feature to obtain a fused feature, wherein the conditional latent is generated based on the fused features (Wang, Section 4.4: “Instead of feeding the scene and the language feature into the self-attention layer, we directly concatenate them as the global conditional feature. We denote this model as w/o self-att. We train and test on the walk subset.”) Claim 10 Examiner’s Interpretation: Machine readable media can encompass forms of signal transmission media that falls outside of the four statutory categories of invention. MPEP 2106; citing In re Nuijten, 500 F.3d 1346, 84 USPQ2d 1495 (Fed. Cir. 2007). A claim whose BRI covers both statutory and non-statutory embodiments embraces subject matter that is not eligible for patent protection and therefore is directed to non-statutory subject matter. MPEP 2106. Claims 10-18 as drafted recite One or more non-transitory computer-readable storage media. The broadest reasonable interpretation of the claimed medium in view of Applicant’s specification covers only eligible subject matter. Claim Mapping: The same teachings and rationales in claim 1 are appliable to claim 10. Claim 11 The same teachings and rationales in claim 2 are appliable to claim 11. Claim 12 The same teachings and rationales in claim 3 are appliable to claim 12. Claim 13 The same teachings and rationales in claim 4 are appliable to claim 13. Claim 17 The same teachings and rationales in claim 8 are appliable to claim 17. Claim 19 The same teachings and rationales in claim 1 are appliable to claim 19, with a CPU/GPU based system disclosing a corresponding system as claimed. Claim(s) 7, 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang, HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes in view of Han, A Survey on Vision Transformer and Choe, PointMixer: MLP-Mixer for Point Cloud Understanding Claim 7 Wang does not explicitly disclose but Choe discloses wherein the down sampling is performed using a k-nearest neighbor classifier (Choe, Section 2: “this paper adopts k-Nearest Neighbor (kNN) for local neighborhood sampling…While downsampling layers adopt pooling with kNN and FPS, upsampling layers re-compute kNN”) Before the effective filing date of this application, it would have been obvious to one of ordinary skill in the art to downsample with KNN. Wang already utilizes downsampling of the point representation, and one of ordinary skill in the art would have motivation to apply the operation to point samples (“or various 3D perception tasks such as object shape classification [80], semantic segmentation [2] and point cloud reconstruction tasks”)(Choe, Section2). One of ordinary skill in the art would have had a reasonable expectation of success because both references consider application to point sampling. Claim 16 The same teachings and rationales in claim 7 are appliable to claim 16. Allowable Subject Matter Claim(s) 5, 6, 9, 14, 15, 18, 20 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim(s) 5, 14, 20 Han, cited to disclose U-net does not suggest application of U-net to “obtaining the pre-trained U-Net scene encoder by pre-training the U-Net scene encoder until a distance between the image feature vectors and the point feature vector is a minimum” Regarding claim(s) 6, 15, Choe, cited to disclose Knn, does not suggest applying an average pooling operation on the set of k-nearest neighboring vectors around each point feature vector of the set of point feature vectors to obtain a plurality of average pooled vectors,wherein the second scene features include the plurality of average pooled vectors Regarding claim(s) 9, 18 Wang considers regularlized losses but uses position rather than size and does not disclose two regularization losses associated with a category of the goal object and a size of the goal object. Additional Prior Art Additional prior art relevant to Applicant’s disclosure but not relied upon: Yi (US 2025/0232506) also discloses a goal state: PNG media_image1.png 487 610 media_image1.png Greyscale Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN M GRAY whose telephone number is (571)272-4582. The examiner can normally be reached on Monday through Friday, 9:00am-5:30pm (EST). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached on (571)272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RYAN M GRAY/Primary Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Feb 10, 2025
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737956
SCENE-AWARE SYNTHETIC HUMAN MOTION GENERATION USING NEURAL NETWORKS
2y 8m to grant Granted Sep 15, 2026
Patent 12737988
GRAPHICAL USER INTERFACE FOR PRESENTING GEOGRAPHIC BOUNDARY ESTIMATION
2y 0m to grant Granted Sep 15, 2026
Patent 12711717
SYSTEM AND METHOD FOR SELECTING TARGETS IN AN AUGMENTED REALITY ENVIRONMENT
2y 2m to grant Granted Aug 18, 2026
Patent 12700169
3D TARGET POINT RENDERING METHOD AND APPARATUS, DEVICE, AND STORAGE MEDIUM
2y 1m to grant Granted Aug 04, 2026
Patent 12694599
APPARENT FORCE CONTROL DEVICE, APPARENT FORCE PRESENTATION SYSTEM, APPARENT FORCE CONTROL METHOD, AND PROGRAM
2y 1m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+11.8%)
2y 0m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 684 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month