Prosecution Insights
Last updated: August 17, 2026
Application No. 18/386,265

DATA AUGMENTATION APPARATUS AND METHOD FOR ACTION RECOGNITION THROUGH SELF-SUPERVISED LEARNING BASED ON OBJECT

Non-Final OA §103§112
Filed
Nov 02, 2023
Priority
Dec 26, 2022 — RE 10-2022-0184018
Examiner
KUDO, KEN
Art Unit
2671
Tech Center
2600 — Communications
Assignee
Gwangju Institute of Science and Technology
OA Round
3 (Non-Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
38 currently pending
Career history
35
Total Applications
across all art units

Statute-Specific Performance

§101
16.1%
-23.9% vs TC avg
§103
51.6%
+11.6% vs TC avg
§102
8.1%
-31.9% vs TC avg
§112
23.4%
-16.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/27/2026 has been entered. Response to Amendment The Amendment filed on May 27th, 2026 has been entered: the application has pending claims 1, 5 and 9. Response to Arguments Applicant’s arguments, see pages 8-17, filed May 27th, 2026, with respect to the rejection(s) of claims 1, 5 and 9 under 35 U.S.C. §112(b) and §103, which was applied to the claim set in the Final Action on 03/09/2026 have been fully considered and are persuasive. Therefore, the previous 35 U.S.C. §112(b) and §103 has been withdrawn following Applicant’s amendment to the claims. However, upon further consideration, a new ground of prior art rejection is made by Yang in view of Shen, further in view of Gao. It is noted the amendments do raise new 112 issues as fully disclosed below. Claim Objections Claims 1 and 5 are objected to because of the following informalities: Claims 1 and 5 recite “a first motion data (x1)”, “a second motion data (x2)”, and “a third motion data (x3)”. The term “data” is a plural or mass noun and is not a countable noun; therefore, using the singular article “a” directly before “data” is grammatically incorrect. Appropriate correction is required. Better wording would be: “first motion data (x1)”, “second motion data (x2)”, “third motion data (x3)”, “new motion data (x̃)”; or amend these phrases to clearly identify a countable object (e.g., “a first set of motion data (x1),…”). Claim 5 is objected to because of the following informalities: Claim 5 introduces sub-steps using the phrases “wherein the extracting the object information feature vector (s) comprises:” and “wherein the generating the new motion data (x̃) is performed by a generator (G) and a discriminator (D) and comprises:”. These constructions are grammatically awkward and can cause confusion as to whether a step or a result is being recited. Appropriate correction is required. Applicant is required to revise these to clearly recite method steps, e.g., “wherein extracting the object information feature vector (s) comprises:” or “wherein the step of extracting the object information feature vector (s) comprises:”, and similarly “wherein generating the new motion data (x̃) is performed by a generator (G) and a discriminator (D) and comprises:”, or equivalent language. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: the terms “an image input unit configured to …”, “an information extraction unit configured to extract ...”, and “a motion information synthesis unit configured to generate ...” in claim 1. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112(a) The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1, 5 and 9 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Independent claims 1 and 5 have been amended to recite a specific machine-learning framework in which: an encoder Es extracts object-information feature vectors s1, s2, and s3 from first, second, and third motion data x1, x2, and x3; the feature vectors are expressly required to represent physical characteristics of the object, including body shape, body height, and posture; second motion data x2 is obtained from first motion data x1 by applying rotation and/or parallel movement, while maintaining the same object information as x1; third motion data x3 is motion data of a different object having different object information from x1; distance-based learning is performed such that a distance between s1 and s2 is decreased and a distance between s1 and s3 is increased, so that encoder Es is trained to extract the claimed object-information feature vector representing the physical characteristics of the object; generator G receives x1 and s1 to regenerate x1; generator G receives x1 and s3 to generate new motion data x̃ in which object information of x1 is substituted with object information of the different object while action information of x1 is maintained; generator G receives x̃ and s1 to regenerate x1; and discriminator D classifies the generated new motion data as fake data while generator G is trained to cause discriminator D to misclassify the new motion data as real data. Although the originally filed disclosure mentions object information such as body shape, body height, and posture, the rejection is not based on the mere presence of those words. Rather, the rejection is based on the absence of adequate written-description support for the specific amended requirement that those physical characteristics are represented by the encoder-produced feature vectors and are learned through the claimed same-object/ different-object distance-based training procedure, then used in the claimed reconstruction/ substitution/ regeneration adversarial framework. The originally filed disclosure does NOT reasonably convey possession of this specific arrangement. At most, the originally filed disclosure describes, at a high level, extracting object-related information and synthesizing new motion data. The disclosure does NOT adequately describe the now-claimed technical relationship among: (i) the physical-characteristic set of body shape, body height, and posture; (ii) the object-information feature vectors s1, s2, and s3; (iii) the deformation-based same-object positive sample x2; (iv) the different-object negative sample x3; (v) the distance-based training of encoder Es; and (vi) the generator/discriminator reconstruction, substitution, regeneration, and adversarial learning sequence. More specifically, the originally filed disclosure fails to provide adequate written-description support for the following amended limitations: Feature vector specifically representing body shape, body height, and posture: the claims now require the object-information feature vector s to represent physical characteristics of the object, including body shape, body height, and posture. The originally filed disclosure does not reasonably convey that encoder Es is trained to output a feature vector specifically representing this claimed set of physical characteristics. A general statement that object information may include physical characteristics does not reasonably convey possession of the more specific claimed arrangement in which the encoder-produced feature vectors s1, s2, and s3 are trained through the claimed distance-based learning to represent body shape, body height, and posture. Distance-based learning tied to those physical characteristics: the claims now require distance-based learning in which the distance between s1 and s2 is decreased and the distance between s1 and s3 is increased, so that encoder Es is trained to extract the object-information feature vector representing the claimed physical characteristics. The originally filed disclosure does not adequately describe this specific training relationship. In particular, the disclosure does not reasonably convey that the same-object/different-object distance relationship is used to train a feature vector specifically representing body shape, body height, and posture, rather than merely describing a desired result of distinguishing object-related information. Deformation-based same-object positive sample: the claims now require that second motion data x2 is obtained by applying a predetermined deformation, including rotation and/or parallel movement, to first motion data x1, with x2 having the same object information as x1. The originally filed disclosure does not adequately describe this claimed deformation-based positive-pair construction as part of a training framework for learning a physical-characteristic feature vector representing body shape, body height, and posture. Substitution of object information while maintaining action information: the claims now require generator G to generate new motion data x̃ in which object information of first motion data x1 is substituted with object information of a different object while action information of x1 is maintained. The originally filed disclosure does not reasonably convey possession of this claimed object/action disentanglement and substitution operation with the specificity now recited. Merely describing generation of new motion data does not provide written-description support for the claimed separation, substitution, and preservation of object information and action information. Reconstruction / substitution / regeneration loop: the claims now require a three-step generator sequence: (i) regenerate x1 from x1 and s1; (ii) generate x̃ from x1 and s3; and (iii) regenerate x1 from x̃ and s1. The originally filed disclosure does not reasonably convey this specific cyclic reconstruction framework as the applicant’s invention. The disclosure does not adequately describe the claimed use of the same physical-characteristic feature vector framework in both the substitution operation and the regeneration operation. Adversarial learning tied to the claimed generated motion data: the claims further require discriminator D to classify the generated new motion data x̃ as fake data and require generator G to cause discriminator D to misclassify x̃ as real data, with performance of generator G improved through repeated adversarial learning. The originally filed disclosure does not adequately describe this adversarial training as applied to the specifically claimed x1/x2/x3, s1/s2/s3, physical-characteristic, substitution, and regeneration framework. The amended claims do not merely restate what was originally disclosed; they combine previously separate concepts into a new operative whole. In the original filing, the applicant described object information extraction, deformation examples, generator-based synthesis, and adversarial learning in a more general and modular fashion. In the RCE, those pieces are now locked into a particular training pipeline with specific sample roles, specific vector roles, and a specific generator/ discriminator loop that was not presented in the original disclosure as a single invented combination. As a result, the amended claims extend beyond the scope of the original written description and introduce new matter under the guise of amendment. The specification, as originally filed, does not provide the clear possession required to support the now-claimed integrated scheme Accordingly, the originally filed disclosure does not reasonably convey possession of the specific amended invention as now claimed. Claims 1 and 5 therefore fail to comply with the written description requirement of 35 U.S.C. §112(a). Claim 9 is rejected for the same reasons because claim 9 recites a non-transitory computer-readable recording medium storing a computer program for executing the method of claim 5, the apparatus of claim 1 and therefore incorporates the unsupported limitations of claims 1 and 5. Claim Rejections - 35 USC § 112(b) The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1, 5 and 9 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites limitations “an image input unit configured to …”, “an information extraction unit configured to extract ...”, and “a motion information synthesis unit configured to generate ...” invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The disclosure in the specification do not provide adequate structure corresponding to the full scope of the claimed functions. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Accordingly, claim 1 fails to particularly point out and distinctly claim the invention because the scope of the §112(f) limitations cannot be determined from the claim language and the specification. Claims 5 and 9 are directed to the method and a non-transitory computer-readable recording medium correspond to the method of Claim 1 , and performs the steps disclosed herein. Therefore, the claims are all rejected. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 5, and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Yang (Yang et al, TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting. arXiv.Org, 2020), in view of Shen (Shen et al. The Imaginative Generative Adversarial Network: Automatic Data Augmentation for Dynamic Skeleton-Based Hand Gesture and Human Action Recognition. arXiv.Org, 2021), further in view of Gao (Gao et al. Contrastive Self-Supervised Learning for Skeleton Action Recognition [Review of Contrastive Self-Supervised Learning for Skeleton Action Recognition]. Proceedings of Machine Learning Research, 148, 51–61, 2021). Regarding claim 1, Yang teaches a data augmentation apparatus for action recognition through self-supervised learning based on objects, comprising: an image input unit configured to input image information; ( [Sec. 3], [Fig. 2]: Yang teaches a motion-retargeting pipeline that receives source and target video inputs, extracts 2D body-joint skeleton sequences from the videos using an off-the-shelf keypoint detector, and uses the extracted skeleton sequences as the input/output space for motion retargeting, thereby teaching inputting image/video information and deriving motion-related skeletal information therefrom. ) an information extraction unit configured to extract an object information feature vector (s) representing object information of an object from motion data (x) obtained from the image information, the object information representing physical characteristics of the object, the physical characteristics including a body shape ( [Abstract], [Sec. 3], [Fig. 2]: Yang teaches extracting 2D body-joint skeleton sequences from source and target videos and using the extracted skeleton sequences as motion data x. Yang further teaches a structure encoder Es(x)=s configured to extract a structure code s from the skeleton/ motion data, where the structure code represents body shape/ body structure of the actor; therby extracting an object-information feature vector representing at least body shape/ body structure from motion data obtained from image/ video information. ) a motion information synthesis unit configured to generate, using the object information feature vector (s), new motion data (x̃) for motion retargeting and animation, ( [Sec. 3], [Fig. 2]: Yang teaches an unsupervised video motion-retargeting framework in which image/ video information is converted into skeleton/ joint motion sequences, and the motion-retargeting network decomposes an input joint sequence into a motion code, a structure code s, and a view code. Yang teaches that the structure code represents body shape/ body structure. Yang further teaches a decoder/ generator G that receives the motion, structure, and view codes and generates/ reconstructs a joint sequence. Yang also teaches combining motion information from a source sequence with structure/ body-shape information from a target subject to generate retargeted skeleton/ motion data having the target body structure while preserving the source motion; therby generating new motion data using an object/ structure information feature vector s for motion retargeting. ) wherein the information extraction unit is configured to: sample a first motion data (x1) of the object; ( [Sec. 3–4], [Fig. 2]: Yang uses motion sequences of a source character represented as pose/ skeleton trajectories as inputs to the retargeting network. Each such sequence corresponds to motion data of one specific character (an “object”); thereby sampling a first motion sequence x1 for a given character. ) obtain a second motion data (x2) by applying a predetermined deformation to the first motion data (x1), the predetermined deformation comprising at least one of rotation (x2) having the same object information as the first motion data (x1); ( [Sec. 3–4], [Figs. 2, 7]: Yang teaches view perturbation of skeleton motion data by rotating a reconstructed 3D joint sequence and projecting the rotated sequence to a 2D joint sequence. The resulting view-perturbed skeleton sequence corresponds to second motion data obtained by applying a predetermined deformation, including rotation, to the first motion data. Because the view-angle perturbation changes the viewpoint/ orientation of the skeleton sequence but corresponds to the same person/ object, the resulting second motion data has the same object/ body-structure information as the first motion data. ) sample a third motion data (x3) of a different object, the third motion data (x3) having different object information from the first motion data (x1); ( [Sec. 3–4]: Yang teaches source and target skeleton sequences from different persons/ objects for motion retargeting, and teaches transferring source motion to a target subject having a different body structure/ body shape. The target/ different-person skeleton sequence corresponds to third motion data x3 of a different object having different object/structure information from the first motion data. ) extract, via an encoder (Es), a first object information feature vector (s1) from the first motion data (x1), a second object information feature vector (s2) from the second motion data (x2), and a third object information feature vector (s3) from the third motion data (x3); and ( [Sec. 3–4]: Yang teaches a structure encoder Es(x)=s that extracts a structure code s from an input skeleton/ motion sequence x. Yang explains that the structure code represents body shape/ body structure and that extracting the structure code can be interpreted as body-shape estimation from the sequence. Therefore, applying Yang’s structure encoder Es to the first, transformed/ rotated second, and different-object third skeleton sequences teaches extracting corresponding structure/ object-information feature vectors s1, s2, and s3 [Yang does not necessarily label them as s1, s2, and s3, but the operation is the same: Es extracts a structure code from each skeleton/ motion sequence]. ) perform distance-based learning over latent representations of motion, structure, and view-angle such that latent codes corresponding to the same underlying factor under different perturbations are encouraged to be similar, and latent codes corresponding to different factors or different instances are encouraged to be dissimilar, so that the encoder (Es) is trained to extract the object information feature vector (s) representing the physical characteristics of the object, ( [Secs. 3.1–3.2.2], [Fig. 5]: Yang’s TransMoMo framework disentangles a structure code representing the physical characteristics (body structure/identity) of a person, separate from motion and view-angle. Yang defines invariance-driven loss functions, including a triplet-style loss used to smooth the structure representation, that operate on the structure codes so that representations for the same person are encouraged to be close while those for different persons are separated in the latent space. This constitutes distance-based learning on the structure feature vectors, decreasing the distance between structure vectors corresponding to the same object and increasing their distance from structure vectors of different objects, thereby training the encoder Es to extract an object-information feature vector s that captures the physical characteristics of the person. ) wherein the motion information synthesis unit comprises a generator (G) and a discriminator (D), ( [Sec. 3], [Fig. 2]: Yang teaches a motion-retargeting network including a decoder/ generator G and a discriminator D. Yang discloses that the decoder/ generator G receives latent motion, structure/ body-shape, and view codes and reconstructs/ generates a 3D joint sequence. Yang further discloses a discriminator D, implemented as a temporal convolutional network. ) wherein the generator (G) is configured to: receive the first motion data (x1) and the first object information feature vector (s1) and regenerate the first motion data (x1); ( [Sec. 3], [Fig. 2]: Yang trains an auto-encoder where the encoder extracts disentangled latent codes for motion, structure (body shape), and view from source and target video clips, and a decoder (generator) G takes combinations of these latent codes and produces a reconstructed 3D joint sequence. When G receives the motion and structure codes derived from a given source sequence, it decodes them to reconstruct that sequence’s 3D joint motion, effectively regenerating the original motion data from the corresponding latent motion and structure representations. ) receive the first motion data (x1) and the third object information feature vector (s3) and generate the new motion data (x̃) in which the object information of the first motion data (x1) is substituted with object information of the different object while action information of the first motion data (x1) is maintained; and ( [Sec. 3, 4.2], [Fig. 2]: Yang’s TransMoMo framework encodes each video into latent motion, structure (body-shape), and view-angle codes, and the decoder/generator G takes arbitrary combinations of these latent codes to produce a 3D joint sequence. In particular, Yang discloses using the motion code extracted from a source person together with the structure code of a different target person to synthesize a retargeted skeleton sequence that preserves the source motion while changing the underlying body structure. Accordingly, Yang teaches using a generator G that receives motion data from a first person and object/structure information of a different person to generate new motion data x̃ where the object information of x1 is substituted by that of the different person while the action (motion) information of x1 is maintained. ) receive the new motion data (x̃) and the first object information feature vector (s1) and regenerate the first motion data (x1), ( [Sec. 3], [Fig. 2]: Yang’s motion-retargeting network encodes each joint sequence into disentangled motion and structure (body-shape) codes and uses a decoder/generator G that takes combinations of these latent codes to produce 3D joint sequences. When the motion code corresponding to a given action is combined with the structure code of the original actor, G reconstructs that actor’s motion; when the same motion code is combined with the structure code of a different actor, G produces a retargeted motion sequence with the same action but different body structure. Therefore, given a retargeted motion sequence x̃ (carrying the source motion) and the original actor’s structure feature s1, Yang’s framework enables the decoder to synthesize the source actor’s motion again, effectively regenerating the first motion data x1 from x̃ and s1. ) wherein the discriminator (D) is configured to classify the new motion data (x̃) generated by the generator (G) as fake data, ( [Sec. 3], [Fig. 2]: Yang teaches a motion-retargeting network including a decoder/ generator G and a discriminator D. Yang’s decoder/ generator G receives latent motion, structure/ body-shape, and view codes and generates/ reconstructs joint/ motion sequences. Yang further teaches using discriminator D, implemented as a temporal convolutional network, in an adversarial framework to distinguish generated/ reconstructed joint sequences from real joint sequences. Therefore, Yang teaches a discriminator configured to classify motion sequences generated by the generator as fake/ generated data during adversarial training. ) wherein the generator (G) is further configured to cause the discriminator (D) to misclassify the new motion data (x̃) as real data, and ( [Sec. 3], [Fig. 2]: Yang’s TransMoMo framework includes an adversarial loss applied on generated joint sequences, where a discriminator D is trained to distinguish real joint sequences from those produced by the generator, and the generator G is trained to produce sequences that are indistinguishable from real data. In this adversarial setup, G is optimized to fool D so that D classifies the generated motion sequences as real rather than fake, which corresponds to configuring the generator to cause the discriminator to misclassify the generated motion data x̃ as real data. ) wherein performance of the generator (G) is improved as adversarial learning between the generator (G) and the discriminator (D) is repeatedly performed. ( [Sec. 3], [Fig. 2]: Yang teaches a motion-retargeting network including a decoder/generator G and a discriminator D. Yang teaches that the decoder/generator G generates/reconstructs joint/motion sequences from latent motion, structure/body-shape, and view codes, while discriminator D is used in an adversarial training framework to distinguish generated/reconstructed joint sequences from real joint sequences. By training the generator and discriminator adversarially, Yang teaches improving the generator/decoder so that the generated skeleton/motion sequences better resemble real skeleton/motion sequences; thereby improving generator performance through adversarial learning between generator G and discriminator D. ) While Yang’s TransMoMo framework focuses on disentangling motion, structure, and view-angle codes for motion retargeting and animation, it does not use the generated motions to augment action-recognition training data. Shen, by contrast, discloses: the object information representing physical characteristics of the object, the physical characteristics including a body height, and a posture of the object; and ( [Abstract], [Secs. 1–3]: Shen teaches skeleton-based hand gesture and human action recognition using skeleton motion data, and explains that recognition performance is affected by physical variations including differences in hand/ human size [corresponds to body height] and hand/ human posture. Shen further teaches that skeleton data differs in size and postures, and that such variations are relevant to data augmentation. Shen also teaches learning latent physical attributes, including human/ hand sizes, and applying those learned attributes to generate synthetic skeleton sequences. ) a motion information synthesis unit configured to generate new motion data (x̃) to augment action recognition training data, ( [Abstract], [Secs. 1–3]: Shen teaches an Imaginative Generative Adversarial Network for automatic data augmentation of dynamic skeleton-based hand gesture and human action recognition. Shen teaches learning latent behavioral and physical attributes from skeleton motion sequences, including physical attributes such as human/ hand sizes, and applying the learned latent attributes to other skeleton data to generate synthetic skeleton sequences. Shen further teaches that the generated synthetic skeleton data is used to augment training data and improve recognition/ classification performance. ) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Yang to use its generated retargeted skeleton/ motion sequences as action-recognition training augmentation, as taught by Shen, because Shen teaches that skeleton-based action recognition benefits from synthetic skeleton data that increases training samples and accounts for human-size and posture variation. The modification would have predictably applied Yang’s generated skeleton sequences to Shen’s known augmentation purpose to improve recognition robustness and reduce overfitting. Yang [as modified by Shen] teaches generating augmented skeleton/ motion data for action recognition, but does not expressly teach the claimed rotation/ parallel-movement positive-pair construction and distance-based feature learning where Gao disclose: obtain a second motion data (x2) by applying a predetermined deformation to the first motion data (x1), the predetermined deformation comprising at least one of rotation and parallel movement, the second motion data (x2) having the same object information as the first motion data (x1); ( [Secs. 3.1–3.2], [Fig. 1]: Gao teaches generating augmented skeleton motion sequences from an original skeleton sequence for contrastive self-supervised learning. Specifically, Gao applies two stochastic transformation setups to the same skeleton sequence to obtain a correlated/ positive pair, including viewpoint and distance transformations. Gao’s transformations include rotation around coordinate axes [rotation] and translation of the observation coordinate system [parallel movement]. These transformations change the viewpoint/ position of the skeleton sequence while preserving the underlying action semantics and the same skeleton/ person information because the transformed sequence is generated from the same original skeleton sequence. Therefore, Gao teaches obtaining a second motion data x2 by applying a predetermined deformation, including rotation and translation/ parallel movement, to a first motion data x1, where x1 and x2 have the same underlying object/body-structure information. ) perform distance-based learning such that a distance between the first object information feature vector (s1) and the second object information feature vector (s2) is decreased, and a distance between the first object information feature vector (s1) and the third object information feature vector (s3) is increased, so that the encoder (Es) is trained to extract an action-representation feature vector (s) invariant motion semantics of the skeleton sequence, ( [Secs. 3.1–3.2], [Fig. 1]: Gao teaches contrastive self-supervised learning for skeleton action recognition in which two transformed views of the same skeleton sequence are generated using viewpoint and distance transformations to form a positive/correlated pair. Gao teaches extracting feature representations from the transformed skeleton sequences using an encoder and applying contrastive loss to maximize agreement between positive pairs while minimizing agreement with negative samples. Thus, Gao teaches distance-based learning in which action-representation feature vectors from transformed views of the same skeleton sequence are made closer, while feature vectors from different skeleton sequences/negative samples are made farther apart, thereby training the encoder to extract a representation capturing invariant motion semantics. ) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Yang [as modified by Shen] by incorporating Gao’s rotation/ translation-based contrastive learning for skeleton sequences. Yang [as modified by Shen] uses skeleton representations and seeks invariant disentanglement of motion, body structure, and view, uses generated skeleton data to improve action-recognition training. Gao teaches that applying rotation and translation transformations to the same skeleton sequence to form positive pairs, and training with contrastive loss against negative samples, improves invariant skeleton representations for action recognition. A POSITA would have been motivated to apply Gao’s known positive-pair contrastive training to Yang [as modified by Shen]’s skeleton feature encoder to improve robustness to viewpoint and position variation in the augmented skeleton data used for action-recognition training, with predictable results because all references operate on skeleton motion sequences. Regarding claims 5 and 9, the rationale in the rejection of claim 1 is provided herein. In addition, the data augmentation apparatus of claim 1 corresponds to the method of claim 5, as well as the non-transitory computer-readable recording medium of claim 9, and performs the steps disclosed herein. Therefore, the claims are all rejected. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. KEN KUDO Examiner Art Unit 2671 /KEN KUDO/Examiner, Art Unit 2671 /VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

Nov 02, 2023
Application Filed
Nov 20, 2025
Non-Final Rejection mailed — §103, §112
Feb 02, 2026
Response Filed
Mar 09, 2026
Final Rejection mailed — §103, §112
May 27, 2026
Request for Continued Examination
Jun 01, 2026
Response after Non-Final Action
Jul 23, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
High
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month