Prosecution Insights
Last updated: October 01, 2026
Application No. 18/639,873

METHOD OF TRAINING A NEURAL NETWORK FOR POSE DETECTION

Final Rejection §103
Filed
Apr 18, 2024
Priority
Mar 04, 2024 — provisional 63/561,204
Examiner
ALAVI, AMIR
Art Unit
2668
Tech Center
2600 — Communications
Assignee
Honda Motor Co., Ltd.
OA Round
2 (Final)
94%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 94% — above average
94%
Career Allowance Rate
1108 granted / 1184 resolved
+31.6% vs TC avg
Minimal +4% lift
Without
With
+3.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
14 currently pending
Career history
1191
Total Applications
across all art units

Statute-Specific Performance

§101
26.9%
-13.1% vs TC avg
§103
20.6%
-19.4% vs TC avg
§102
19.7%
-20.3% vs TC avg
§112
8.7%
-31.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1184 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s amendment filed on 13, July 2026, has been entered and made of record. Applicant’s arguments with respect to claims 1-10 have been considered but are moot in view of the new amendments to independent claim 1. Applicant argues in essence that the cited Prior Arts, do not address, “identifying a plurality of action segments in the test video and a duration associated with each action segment using the trained video encoder.”. Examiner respectfully disagrees and indicates that the cited Prior Art reasonably address limitations of the claimed invention. Applicant is reminded that Examiner will interpret each claim in the broadest reasonable sense, as such, the claims and only the claims form the metes and bounds of the invention. In this regard, Examiner directs your attention to Nie et al. (CN 114596520 A), page 9, 10th paragraph, “interframe coding module Inter-frame Encoder, used for modeling time process, using self-attention mechanism to finish the processing of action time sequence information the long-term relation of modeling movement, rationally distributing the weight of the inter-frame feature embedding, and can inhibit the interference of the irrelevant target and object in the video, distributing more attention resources for the focus area”. Nonetheless, Examiner looks forward to receiving further amendments/Interviews and to expedite allowance of this application. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-10 are rejected under 35 U.S.C. 103(a) as being unpatentable over Nie et al. (CN 114596520 A, A First Visual Angle Video Action Recognition Method And Device), hereinafter, “Nie”, in view of Cai et al. (CN 115909406 A, A Gesture Recognition Method Based On Multi-class Classification), hereinafter, “Cai”, and further in view of Guo et al. (USPAP 2025/0378,806), hereinafter, “Guo”. Regarding claim 1 Nie teaches, obtaining a test video with a plurality of actions (Please note, page 3, 4th paragraph. As indicated a first visual angle video action recognition the first visual angle video action data set input based on RGB mode and depth mode of multi-scale network extracting space semanteme.); inputting RGB features from the test video to a video encoder (Please note, page 3, 4th paragraph. As indicated to obtain the feature embedding vector with rich multi-scale bimodal space semantics, processing the feature embedding vector as the input of Inter-frame Encoder module.); applying the video encoder to the RGB features from the test video to output RGB feature embeddings. (Please note, page 3, 4th paragraph. As indicated through multiple Inter-frame processing the Encoder module, finishing the extraction of the inter-frame timing relationship, obtaining three feature embedding respectively, depth branch and multi-scale fusion branch, fusing the data of the RGB branch and the depth branch through the CFAM module, and finishing the fusion of the feature embedding vector of the multi-scale fusion branch, generating combined feature embedding vector, processing the combined feature embedding vector through the linear layer, obtaining the action classification result of each frame, then averaging the video frame of one action segment along the time sequence direction, outputting the final result of the recognition.); identifying a plurality of action segments in the test video and a duration associated with each action segment using the trained video encoder. (Please note, page 9, 10th paragraph. As indicated interframe coding module Inter-frame Encoder, used for modeling time process, using self-attention mechanism to finish the processing of action time sequence information the long-term relation of modeling movement, rationally distributing the weight of the inter-frame feature embedding, and can inhibit the interference of the irrelevant target and object in the video, distributing more attention resources for the focus area.). Nie does not recite, inputting pose features from the test video to a pose encoder; applying the pose encoder to the pose features from the test video to output pose feature embeddings; mapping the RGB feature embeddings and the pose feature embeddings to a shared representation space Cai teaches, inputting pose features from the test video to a pose encoder; applying the pose encoder to the pose features from the test video to output pose feature embeddings (Please note, page 3, 1st paragraph. As indicated inputting the hand RGB image into the 2 D hand posture estimation network and the hand segmentation network, extracting the 2 D hand posture primary feature by the encoder in the 2 D hand posture estimation network, extracting the hand segmentation primary feature by the encoder in the hand segmentation network, the 2 D hand posture primary feature and hand divided primary feature input information sharing module, respectively obtaining the fusion 2 D hand posture mid-stage feature and fusion hand dividing stage feature.); mapping the RGB feature embeddings and the pose feature embeddings to a shared representation space. (Please note, page 3, 1st paragraph .As indicated the fusion 2 D hand posture middle-stage feature returns to the 2 D hand gesture estimation network branch, performing residual fusion with the 2 D hand posture primary characteristic to obtain the 2 D hand posture advanced characteristic, returning the fusion hand segmentation middle level characteristic to the hand segmentation network branch, performing residual fusion with the hand segmentation primary characteristic to obtain the hand segmentation high-level characteristic.) Nie & Cai are combinable because they are from the same field of endeavor. At the time before the effective filing date, it would have been obvious to a person of ordinary skill in the art to utilize this inputting pose features from the test video to a pose encoder; applying the pose encoder to the pose features from the test video to output pose feature embeddings; mapping the RGB feature embeddings and the pose feature embeddings to a shared representation space of Cai in Nie’s invention. The suggestion/motivation for doing so would have been as indicated on page 3, 1st paragraph, “to obtain the hand segmentation high-level characteristic”. Nie and Cai do not expressly recite, utilizing a contrastive loss to train the video encoder; and identifying one or more action segments in the test video using the trained video encoder.. Guo recites utilizing a contrastive loss to train the video encoder (Please note, paragraph 0064. As indicated determining a contrastive loss based on the training video feature and the training music feature and jointly training the video encoder and the music encoder based on the contrastive loss.); and identifying one or more action segments in the test video using the trained video encoder. (Please note, paragraph 0006. As indicated an action select box to select an action to be performed by the at least one object, and a trigger select box to select a trigger for triggering the action.). Nie, Cai & Guo are combinable because they are from the same field of endeavor. At the time of the invention, it would have been obvious to a person of ordinary skill in the art to utilize this utilizing a contrastive loss to train the video encoder; and identifying one or more action segments in the test video using the trained video encoder of Guo in Nie & Cai’s invention. The suggestion/motivation for doing so would have been as indicated on paragraph 0037, “The video encoder 314 and the music encoder 324 in the DSSM 300A may be jointly trained based on the contrastive loss 328.”. Therefore, it would have been obvious to combine Nie, Cai with Guo to obtain the invention as specified in claim 1. Regarding claim 2, Guo recites, wherein the contrastive loss is determined using at least one contrastive learning process. (Please note, paragraph 0064. As indicated jointly training the video encoder and the music encoder based on the contrastive loss.). Regarding claim 3, Guo recites, wherein the at least one contrastive learning process includes a vanilla contrastive learning process. (Please note, paragraph 0037. As indicated a contrastive loss 328 may be determined based on the training video feature 316 and the corresponding training music feature 326. The video encoder 314 and the music encoder 324 in the DSSM 300A may be jointly trained based on the contrastive loss 328. In this regard, Examiner considers this basic contrastive to correspond to Applicant’s vanilla contrastive.). Regarding claim 4, Guo recites, wherein the at least one contrastive learning process includes a pose-supervised contrastive learning process. (Please note, paragraph 0043. As indicated the pixel changes between these frames may be quantified, which serves as an indicator of the motion's vigor. Essentially, greater pixel variation suggests more intense motion, while minimal changes imply a slower pace or static scene. In this regard, Examiner considers this static scene to correspond to Applicant’s pose-supervised contrastive.). Regarding claim 5, Guo recites, wherein the pose-supervised contrastive learning process is associated with a predetermined threshold value for determining negative pair sets. (Please note, paragraph 0067. As indicated a correlation level between the variance intensity of the music content and a motion intensity indicated by the motion information is greater than a threshold. In this regard, Examiner considers this, greater than a threshold to correspond to Applicant’s negative pair.). Regarding claim 6, Guo recites, wherein the at least one contrastive learning process includes an action-supervised contrastive learning process. (Please note, paragraph 0051. As indicated The energy variations of both the music and video are encapsulated in numerical arrays, enabling the application of a correlation coefficient to quantify the congruence between the music's energy trajectory and the video's motion intensity profile. Additionally, the magnitude of the music's energy change at the instant of the video's most significant motion intensity shift is calculated. These metrics are integrated to formulate a correlation level that evaluates the structural compatibility and synchronization precision of the music relative to the video's dynamic progression.). Regarding claim 7, Cai recites, wherein the RGB features and the pose features are associated with a plurality of keypoints of a human subject in the test video. (Please note, page 3, 1st paragraph. As indicated respectively inputting the hand RGB image into the 2 D hand posture estimation network and the hand segmentation network, extracting the 2 D hand posture primary feature by the encoder in the 2 D hand posture estimation network.). Regarding claim 8, Cai recites, wherein the plurality of keypoints includes reference points associated with a face, hands, and/or a body of the human subject in the test video. (Please note, page 3, last paragraph. As indicated the 2 D hand posture primary feature, fusion 2 D hand posture feature and the 2 D hand posture advanced feature fusion, the fusion result and the 2 D hand joint thermogram fusion to obtain the 2 D hand posture feature fusion.). Regarding claim 9, Cai recites, wherein the plurality of keypoints are normalized prior to the step of mapping the RGB embeddings and the pose embeddings to the shared representation space. (Please note, page 6, 2nd paragraph. As indicated keeping complete picture for the 2 D hand joint thermogram, concentrating the hand divided area probability map into small scale convolution kernel from the original scale; scanning and filtering the 2 D hand joint thermogram complete picture by the filter made of small scale convolution kernel; performing convolution operation to obtain the fusion 2D hand posture mid-level characteristic.). Regarding claim 10, Cai recites, extracting the pose features from the test video using a pre-trained pose extractor. (Please note, page 5, 4th paragraph. As indicated inputting the hand RGB image into the 2 D hand posture estimation network and the hand segmentation network, extracting the 2 D hand posture primary feature by the encoder in the 2 D hand posture estimation network.). Allowable Subject Matter Claims 21-24 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The closest applied Prior Art of record fails to disclose or reasonably suggest wherein normalizing the plurality of keypoints comprises centering and scaling each keypoint of the plurality of keypoints with respect to a center of mass of the human subject and wherein the action-supervised contrastive learning process includes selecting an anchor frame in the test video and determining whether other frames in the test video include poses having an action label with a same class as an action label associated with the anchor frame. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Examiner’s Note The examiner cites particular figures, paragraphs, columns and line numbers in the references as applied to the claims for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claims, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner. Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMIR ALAVI whose telephone number is (571)272-7386. The examiner can normally be reached on M-F from 8:00-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571)272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMIR ALAVI/Primary Examiner, Art Unit 2668 Wednesday, September 2, 2026
Read full office action

Prosecution Timeline

Apr 18, 2024
Application Filed
Apr 22, 2026
Non-Final Rejection mailed — §103
Jul 13, 2026
Response Filed
Sep 04, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749147
GENERATION SUPER SAMPLING
2y 9m to grant Granted Sep 29, 2026
Patent 12725419
DETECTING APPARATUS, POSITION CALCULATION SYSTEM, AND DETECTING METHOD
2y 5m to grant Granted Sep 01, 2026
Patent 12725450
INFORMATION PROCESSING METHOD, INFORMATION PROCESSING SYSTEM, AND RECORDING MEDIUM
2y 1m to grant Granted Sep 01, 2026
Patent 12718390
ADAPTIVE FRAME RATE FOR LOW POWER SLAM IN XR
2y 6m to grant Granted Aug 25, 2026
Patent 12711808
ROBUST AND LONG-RANGE MULTI-PERSON IDENTIFICATION USING MULTI-TASK LEARNING
2y 7m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
94%
Grant Probability
97%
With Interview (+3.8%)
2y 3m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1184 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month