Prosecution Insights
Last updated: August 18, 2026
Application No. 18/859,200

EVALUATION METHOD, PROGRAM, AND EVALUATION SYSTEM

Non-Final OA §103
Filed
Oct 23, 2024
Priority
May 12, 2022 — JP 2022-079051 +1 more
Examiner
NAH, JONGBONG
Art Unit
Tech Center
Assignee
Omron Corporation
OA Round
1 (Non-Final)
76%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
90 granted / 118 resolved
+16.3% vs TC avg
Strong +16% interview lift
Without
With
+16.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
24 currently pending
Career history
135
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
64.3%
+24.3% vs TC avg
§102
22.3%
-17.7% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 118 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 10/23/2024 and 10/09/2025 is/are compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Office Action Summary Claim(s) 2-3 and 13-14 is/are canceled. Claim(s) 1, 4-5, and 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al (Predicting Pedestrian Crossing Intention With Feature Fusion and Spatio-Temporal Attention; hereinafter “Yang (2021)”) in view of Yang et al (Resolution Adaptive Networks for Efficient Inference; hereinafter “Yang (2020)”), further in view of Yan et al (Robust Multi-Resolution Pedestrian Detection in Traffic Scenes). Claim(s) 6-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al (Predicting Pedestrian Crossing Intention With Feature Fusion and Spatio-Temporal Attention; hereinafter “Yang (2021)”) in view of Yang et al (Resolution Adaptive Networks for Efficient Inference; hereinafter “Yang (2020)”) and Yan et al (Robust Multi-Resolution Pedestrian Detection in Traffic Scenes), further in view of Rasouli et al (Joint Attention in Driver-Pedestrian Interaction: from Theory to Practice). Claim(s) 12 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al (Predicting Pedestrian Crossing Intention With Feature Fusion and Spatio-Temporal Attention; hereinafter “Yang (2021)”) in view of Rehder et al (Head Detection and Orientation Estimation for Pedestrian Safety) and Yan et al (Robust Multi-Resolution Pedestrian Detection in Traffic Scenes), further in view of Yang et al (Resolution Adaptive Networks for Efficient Inference; hereinafter “Yang (2020)”). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 4-5, and 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al (Predicting Pedestrian Crossing Intention With Feature Fusion and Spatio-Temporal Attention; hereinafter “Yang (2021)”) in view of Yang et al (Resolution Adaptive Networks for Efficient Inference; hereinafter “Yang (2020)”), further in view of Yan et al (Robust Multi-Resolution Pedestrian Detection in Traffic Scenes). Regarding claim(s) 1 and 16, Yang (2021) teaches an evaluation method performed by an arithmetic circuit accessible to a storage device storing a plurality of trained models, the plurality of trained models including: (Figure 2; and Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “The task of vision-based pedestrian crossing intention prediction is formulated as follows. Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted. They are used together with the vehicle’s speed as separate channels that serve as the input to the prediction model”); and (read as “pose keypoints”) based on one or plurality of parts of the movable object, output evaluation of the moving direction of the movable object (Figure 2; Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “The task of vision-based pedestrian crossing intention prediction is formulated as follows. Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted. They are used together with the vehicle’s speed as separate channels that serve as the input to the prediction model”; and Page 4, III. B. Input acquisition, Pedestrian pose keypoints: “Pedestrian pose keypoints represent the target pedestrian’s detailed motion, i.e., the posture at each frame while moving […] we utilize pre-trained OpenPose model to extract the pedestrian pose keypoints […] where p is a 36D vector of 2D coordinates that contain 18 pose joints […]”), and the evaluation method comprising: (Figure 2; and Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “The task of vision-based pedestrian crossing intention prediction is formulated as follows. Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted. They are used together with the vehicle’s speed as separate channels that serve as the input to the prediction model”); and (Figure 2; Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “The task of vision-based pedestrian crossing intention prediction is formulated as follows. Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted. They are used together with the vehicle’s speed as separate channels that serve as the input to the prediction model”; and Page 4, III. B. Input acquisition, Pedestrian pose keypoints: “Pedestrian pose keypoints represent the target pedestrian’s detailed motion, i.e., the posture at each frame while moving […] we utilize pre-trained OpenPose model to extract the pedestrian pose keypoints […] where p is a 36D vector of 2D coordinates that contain 18 pose joints […]”). Yang (2021) fails to teach a simplified model trained; and a detailed model trained; when a resolution of an image of the movable object detected from target image information is smaller than a threshold value, selecting the simplified model from the plurality of trained models However, Yang (2020) teaches wherein the plurality of trained models includes: a simplified (read as “lightweight sub-network”) model trained and a detailed (read as “higher resolution sub-networks”) model trained (Figure 1; Figure 2; Abstract: “In RANet, the input images are first routed to a lightweight sub-network that efficiently extracts low-resolution representations, and those samples with high prediction confidence will exit early from the network without being further processed. Meanwhile, high-resolution paths in the network maintain the capability to recognize the “hard” samples”; Page 3, 3.2. Overall Architecture, 1st Paragraph: “Figure 2 illustrates the overall architecture of the proposed RANet. It contains an Initial Layer and H subnetworks corresponding to different resolutions”; and Page 2, Left Col., 3rd Paragraph: “The “easy” samples are classified by the sub-network with the feature maps in the lowest spatial resolution. The sub-networks with higher resolution will be applied when the previous sub-network fails to achieve a given criterion”); when a resolution of an image of the movable object detected from target image information is smaller than a threshold value, selecting the simplified (read as “lightweight sub-network”) model from the plurality of trained models(Figure 1; Figure 2; Abstract: “In RANet, the input images are first routed to a lightweight sub-network that efficiently extracts low-resolution representations, and those samples with high prediction confidence will exit early from the network without being further processed. Meanwhile, high-resolution paths in the network maintain the capability to recognize the “hard” samples”; and Page 2, Left Col., 3rd Paragraph: “The “easy” samples are classified by the sub-network with the feature maps in the lowest spatial resolution. The sub-networks with higher resolution will be applied when the previous sub-network fails to achieve a given criterion”); and when the resolution is equal to or greater than the threshold value, selecting the detailed (read as “higher resolution sub-networks”) model from the plurality of trained models(Figure 1; Figure 2; Abstract: “In RANet, the input images are first routed to a lightweight sub-network that efficiently extracts low-resolution representations, and those samples with high prediction confidence will exit early from the network without being further processed. Meanwhile, high-resolution paths in the network maintain the capability to recognize the “hard” samples”; and Page 2, Left Col., 3rd Paragraph: “The “easy” samples are classified by the sub-network with the feature maps in the lowest spatial resolution. The sub-networks with higher resolution will be applied when the previous sub-network fails to achieve a given criterion”). Yang (2021) disclose a pedestrian intention prediction framework that predicts whether a pedestrian will cross the roadway using multiple feature modalities, including bounding-box trajectory, local context, global context, pose keypoints, and ego-vehicle speed, thereby teaching the use of whole information and part information to generate a movement-related prediction. Yang (2020) teaches a neural-network architecture including multiple trained subnetworks having different computational complexity, wherein input images are first processed by a lightweight sub-network and are subsequently processed by higher-resolution subnetworks only when additional processing is required. Therefore, it would have been obvious to one of ordinary skill in the art to combine before the effective filing date of the claimed invention to modify the pedestrian intention prediction framework of Yang (2021) by incorporating the adaptive multi-subnetwork architecture of Yang (2020) so that pedestrian movement prediction could be performed using different trained models according to the complexity of the input sample, thereby improving computational efficiency while maintaining prediction capability for more difficult samples. The motivation for this combination of references would have been to achieve a better trade-off between accuracy and computational cost by representations for easy samples while maintaining high-resolution processing capability for hard samples, as taught by Yang (2020). This motivation for the combination of Yang (2021) and Yang (2020) is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Yang (2021) and Yang (2020) fail to teach when a resolution of an image of the movable object detected from target image information is smaller than a threshold value; and when the resolution is equal to or greater than the threshold value. However, Yan teaches when a resolution of an image of the movable object detected from target image information is smaller than a threshold value and when the resolution is equal to or greater than the threshold value (Figure 2; Abstract: “The serious performance decline with decreasing resolution is the major bottleneck for current pedestrian detection techniques. In this paper, we take pedestrian detection in different resolutions as different but related problems, and propose a Multi-Task model to jointly consider their commonness and differences”; Page 3035, Left Col., 2nd Paragraph: “[…] we consider the partition of two resolutions (low resolution: 30-80 pixels tall, and high resolution: taller than 80 pixels[…]”; and Page 3034, 3. Multi-Task Deformable Part Model, 1st Paragraph: “There are two intuitive strategies to handle the multiresolution detection. One is to combine samples from different resolutions to train a single detector (Fig. 2(a)), and another is to train independent detectors for different resolutions (Fig. 2(b))”). Yang (2021) disclose a pedestrian intention prediction framework that predicts whether a pedestrian will cross the roadway using multiple feature modalities including bounding-box trajectory, local context, global context, pose keypoints, and ego-vehicle speed, thereby teaching the use of whole information and part information for generating a movement-related prediction. Yang (2021) teaches an adaptive neural-network architecture including multiple trained sub-networks having different computational complexity, wherein a lightweight sub-network initially processes the input and higher-resolution subnetworks are selectively invoked when additional processing is required. Yan further teaches that pedestrians having different image resolutions should be treated differently by categorizing pedestrians into low-resolution and high-resolution groups and employing different resolution-aware detection strategies. Therefore, it would have been obvious to one of ordinary skill in the art to combine before the effective filing date of the claimed invention to modify the pedestrian intention prediction framework of Yang (2020) by incorporating the adaptive multi-subnetwork architecture of Yang (2021) and applying the resolution dependent processing strategy taught by Yan so that pedestrian movement prediction is performed using an appropriate prediction model according to the image resolution of the pedestrian. The motivation for this combination of references would have been to achieve a better trade-off between accuracy and computational cost while recognizing that pedestrian detection in different resolutions should be treated as different but related problems, thereby applying different prediction models according to the image resolution of the pedestrian. This motivation for the combination of Yang (2020), Yang (2021), and Yan is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Regarding claim(s) 4, Yang (2021) as modified by Yang (2020) and Yan teaches the evaluation method according to claim 1, wherein the resolution is determined based on an area of a bounding box of the movable object detected from the target image information (where Yang (2021) teaches in Figure 2; and Page 3, III. A. Problem formulation, 2nd Paragraph: “[…] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted”; and where Yan teaches in Figure 1; Figure 2; Page 3033, 1. Introduction, 1st Paragraph: “[…] however, they encounter difficulties for the low resolution pedestrians (e.g. 30-80 pixels tall, Fig. 1) […]”; Page 3033, Right Col., 2nd Paragraph: “[…] the best detector achieves 21% mean miss rate for pedestrians taller than 80 pixels in Caltech Pedestrian Benchmark [14], while increases to 73% for pedestrians 30-80 pixels high”; and Page 3035, Left Col., 2nd Paragraph: “[…] we consider the partition of two resolutions (low resolution: 30-80 pixels tall, and high resolution: taller than 80 pixels[…]”). Regarding claim(s) 5, Yang (2021) as modified by Yang (2020) and Yan teaches the evaluation method according to claim 1, where Yang (2021) teaches wherein the one or plurality of pieces of whole information include at least one of: a position of the movable object, a speed of the movable object, or an image of a whole of the movable object (Figure 2; Page 3, III. A. Problem formulation, 2nd Paragraph: “[…] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted. They are used together with the vehicle’s speed as separate channels that serve as the input to the prediction model. Therefore, our model has the following input sources: […]”; and Page 4, III. B. Input acquisition, Local context and 2D location trajectory: “2D location trajectory Li give the position change of the target pedestrian in the image […]”). Regarding claim(s) 15, Yang (2021) as modified by Yang (2020) and Yan teaches the a non-transitory storage medium storing a program for performing the evaluation method according to claim 1, where Yang (2021) teaches by the arithmetic circuit (Figure 2; Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “The task of vision-based pedestrian crossing intention prediction is formulated as follows […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted”; and Page 4, C. Model architecture, 1st Paragraph: “The overall architecture […] It consists of CNN modules, RNN modules, attention modules, and a novel way of fusing different features”). Claim(s) 6-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al (Predicting Pedestrian Crossing Intention With Feature Fusion and Spatio-Temporal Attention; hereinafter “Yang (2021)”) in view of Yang et al (Resolution Adaptive Networks for Efficient Inference; hereinafter “Yang (2020)”) and Yan et al (Robust Multi-Resolution Pedestrian Detection in Traffic Scenes), further in view of Rasouli et al (Joint Attention in Driver-Pedestrian Interaction: from Theory to Practice). Regarding claim(s) 6, Yang (2021) as modified by Yang (2020) and Yan teaches the evaluation method according to claim 1, where Yang (2021) teaches wherein the one or plurality of pieces of part information include at least one of: one or more positions of the one or plurality of parts of the movable object (Figure 2; Page 3, III. A. Problem formulation, 2nd Paragraph: “In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted”; and Page 4, III. B. Input acquisition, Pedestrian pose keypoints: “Pedestrian pose keypoints represent the target pedestrian’s detailed motion, i.e., the posture at each frame while moving […] we utilize pre-trained OpenPose model to extract the pedestrian pose keypoints […] where p is a 36D vector of 2D coordinates that contain 18 pose joints […]”). Yang (2021), Yang (2020) and Yan fail to teach one or more directions of the one or plurality of parts of the movable object; one or more images of the one or plurality of parts of the movable object; or information based on one or more relationships among the plurality of parts of the movable object. However, Rasouli teaches one or more directions of the one or plurality of parts of the movable object (Page 53, 7th Paragraph: “[…] The tree structure that describes the overall pose is based on the position of the joints starting from the nose position all the way down to the ankles”; and Page 53, 5th Paragraph: “to learn the interdependencies between the connected parts. Here, the learning is in the form of a spatial prior that describes the relative arrangement of the parts both in terms of the location and orientation of the parts”); one or more images of the one or plurality of parts of the movable object (Page 53, 2nd Paragraph: “a large dataset of greyscale head pose images, each corresponding to a specific head orientation. These images are then normalized and used to train an SVM model to learn the correspondences between the 2D images and 3D head poses”); or information based on one or more relationships among the plurality of parts of the movable object (Page 53, 4th Paragraph: “Part-based models […] consist of two stages: learning human body part appearances, and determining the relationship between them”; Page 53, 5th Paragraph: “to learn the interdependencies between the connected parts. Here, the learning is in the form of a spatial prior that describes the relative arrangement of the parts both in terms of the location and orientation of the parts”). Therefore, it would have been obvious to one of ordinary skill in the art to combine before the effective filing date of the claimed invention to incorporating the additional body-part information taught by Rasouli into Yang's (2021) pedestrian intention prediction framework would have predictably improved the representation of pedestrian posture and behavior, thereby enhancing the accuracy and robustness of pedestrian intention prediction. The motivation for this combination of references would have been to provide rich information about pedestrian behavior and intention by learning human body part appearances and determining the relationship between them. This motivation for the combination of Yang (2021), Yang (2020), Yan and Rasouli is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Regarding claim(s) 7, Yang (2021) as modified by Yang (2020), Yan, and Rasouli teaches the evaluation method according to claim 6, wherein the one or more positions of the one or plurality of parts include a position of a face of the movable object (where Yang (2021) teaches in Page 2, II. Related Work, Feature fusion, 2nd Paragraph: “human poses/skeletons in pedestrian crossing prediction tasks since human pose can be considered as a good indicator of human behaviors. By extracting the pose keypoints from cropped pedestrian images, crossing behavior classifiers are built based on the human pose feature vectors”; and Page 4, III. B. Input acquisition, Pedestrian pose keypoints: “Pedestrian pose keypoints represent the target pedestrian’s detailed motion, i.e., the posture at each frame while moving […] we utilize pre-trained OpenPose model to extract the pedestrian pose keypoints […] where p is a 36D vector of 2D coordinates that contain 18 pose joints […]”; and where Rasouli teaches in Figure 31; Page 32, 9.2.2 Learning, 1st Paragraph: “the instructor's face is first detected using Haar-like features. Next, the 3D pose of the face is estimated using optical flow information based on which a series of hypotheses are generated to estimate the gaze direction of the instructor”; Page 30, 9.2.1 Socially Interactive Systems, 4th Paragraph: “The gaze finding algorithm […] regions that may contain the instructor's face are selected, then a face detection algorithm is run to locate the face of the instructor”; and Page 32, 9.2.2 Learning, 2nd Paragraph: “it then performs a face detection to identify the instructor’s face”). Regarding claim(s) 8, Yang (2021) as modified by Yang (2020), Yan, and Rasouli teaches the evaluation method according to claim 6, where Rasouli teaches wherein the one or more directions of the one or plurality of parts include at least one of: a direction of a face of the movable object, or a line of sight of the movable object (Figure 31; Figure 38; Page 32, 9.2.2 Learning, 1st Paragraph: “the instructor's face is first detected using Haar-like features. Next, the 3D pose of the face is estimated using optical flow information based on which a series of hypotheses are generated to estimate the gaze direction of the instructor. The final gaze is estimated using a tolerable error threshold […]”; and Page 71, 2nd Paragraph: “These algorithms mainly rely on the dynamics of the vehicle and the pedestrian and some also take into account pedestrians' awareness by observing their head orientation”). Regarding claim(s) 9, Yang (2021) as modified by Yang (2020), Yan, and Rasouli teaches the evaluation method according to claim 6, where Rasouli teaches wherein the one or more images of the one or plurality of parts include an image of a face of the movable object (Page 53, 2nd Paragraph: “a large dataset of greyscale head pose images, each corresponding to a specific head orientation. These images are then normalized and used to train an SVM model to learn the correspondences between the 2D images and 3D head poses”; and Page 30, 9.2.1 Socially Interactive Systems, 4th Paragraph: “Once identified, the robot fixates on the face and a second camera with zoom capability captures a close-up image of the face. The close range image allows the robot to fond the instructor's eyes and as a result determine her gaze direction”). Regarding claim(s) 10, Yang (2021) as modified by Yang (2020), Yan, and Rasouli teaches the evaluation method according to claim 6, where Rasouli teaches wherein the information based on one or more relationships among the plurality of parts includes information on a pose of the movable object (Page 53, 4th Paragraph: “Part-based models […] consist of two stages: learning human body part appearances, and determining the relationship between them”; Page 53, 5th Paragraph: “to learn the interdependencies between the connected parts. Here, the learning is in the form of a spatial prior that describes the relative arrangement of the parts both in terms of the location and orientation of the parts”; Page 53, 7th Paragraph: “[…] The tree structure that describes the overall pose is based on the position of the joints starting from the nose position all the way down to the ankles”; and Page 53, 5th Paragraph: “to learn the interdependencies between the connected parts. Here, the learning is in the form of a spatial prior that describes the relative arrangement of the parts both in terms of the location and orientation of the parts”). Regarding claim(s) 11, Yang (2021) as modified by Yang (2020) and Yan teaches the evaluation method according claim 1, where Yang (2021) teaches wherein: evaluation of a moving direction of the movable object includes a probability that the moving direction of the movable object is a first moving direction and a probability that the moving direction of the movable object is a second moving direction (Figure 2; and Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “The task of vision-based pedestrian crossing intention prediction is formulated as follows. Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted. They are used together with the vehicle’s speed as separate channels that serve as the input to the prediction model”). Yang (2021), Yang (2020) and Yan fail to teach where the first moving direction is a direction in which the movable object moves toward a target object; and the second moving direction is a direction in which the movable object avoids the target object. However, Rasouli teaches where the first moving direction is a direction in which the movable object moves toward a target object (Page 59, 11.3.2 Pedestrian Intention Estimation, 5th Paragraph: “Besides dynamic information, these techniques assume a potential goal for pedestrians based on which their trajectories are predicted […] they assign to pedestrians a predefined goal location, which is based on the current orientation of the pedestrian towards one of the potential goal locations”; and Page 50, 11.2 What the Pedestrian […] Pose Estimation and Activity Recognition, 1st Paragraph: “For instance, it helps identifying someone's walking direction towards the street or handwave to send a signal”); and the second moving direction is a direction in which the movable object avoids the target object (Page 63, 3rd Paragraph: “CBS have been used in robotic applications such as navigation and obstacle avoidance and autonomous driving […] the perception module identifies the features (e.g. signs, road users, etc.) in the scene. Then, depending on the criticality of each feature, a value between 0 and 1 is associated with it”; and Page 59, 11.3.2 Pedestrian Intention Estimation, 6th Paragraph: “a graphical model that takes into account factors such as pedestrian trajectory, distance to the curb and awareness [...] the pedestrian looking towards the car is a sign that he noticed the car and is less likely to cross the street. This model, however, is based on a scripted data which means that the participants were instructed to perform certain actions”). Yang (2021) teaches estimating the probability of predicted pedestrian actions and Rasouli teaches predicting pedestrian movement toward a potential goal location based on pedestrian orientation and walking direction, and teaches autonomous-driving navigation including obstacle avoidance based on detected scene features. Therefore, it would have been obvious to one of ordinary skill in the art to combine before the effective filing date of the claimed invention to predictably enabled evaluating alternative movement directions of a pedestrian, including movement toward a destination and movement that avoids an object, while providing probabilistic evaluation of such movement outcomes. The motivation for this combination of references would have been to estimate the probability of pedestrian behavior, predict trajectories toward a potential goal location, and perform autonomous-driving navigation and obstacle avoidance based on perceived scene features. This motivation for the combination of Yang (2021), Yang (2020), Yan and Rasouli is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Claim(s) 12 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al (Predicting Pedestrian Crossing Intention With Feature Fusion and Spatio-Temporal Attention; hereinafter “Yang (2021)”) in view of Rehder et al (Head Detection and Orientation Estimation for Pedestrian Safety) and Yan et al (Robust Multi-Resolution Pedestrian Detection in Traffic Scenes), further in view of Yang et al (Resolution Adaptive Networks for Efficient Inference; hereinafter “Yang (2020)”). Regarding claim(s) 12 and 17, Yang (2021) teaches an evaluation method performed by an arithmetic circuit accessible to a storage device storing a plurality of trained models, the plurality of trained models including: a first model trained to (Figure 2; Page 2, Left Col., 3rd Paragraph: “The proposed method employs a novel neural network architecture for utilizing different spatio-temporal features with a hybrid fusion strategy”; and Page 4, C. Model architecture, 1st Paragraph: “The overall architecture […] It consists of CNN modules, RNN modules, attention modules, and a novel way of fusing different features”), in response to input of one or plurality of pieces of whole information based on a whole of a movable object in image information and one or plurality of pieces of first part information based on one or plurality of first parts of the movable object, output evaluation of a moving direction of the movable object (Figure 2; Page 2, Feature fusion, 2nd Paragraph: “By extracting the pose keypoints from cropped pedestrian images, crossing behavior classifiers are built based on the human pose feature vectors. Improvement in prediction accuracy shows the effectiveness of using pose features”; Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “[…] Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted […]”; and Page 4, III. B. Input acquisition, Pedestrian pose keypoints: “Pedestrian pose keypoints represent the target pedestrian’s detailed motion, i.e., the posture at each frame while moving […] we utilize pre-trained OpenPose model to extract the pedestrian pose keypoints […] where p is a 36D vector of 2D coordinates that contain 18 pose joints […]”); and a second model trained to (Figure 2; Page 2, Left Col., 3rd Paragraph: “The proposed method employs a novel neural network architecture for utilizing different spatio-temporal features with a hybrid fusion strategy”; and Page 4, C. Model architecture, 1st Paragraph: “The overall architecture […] It consists of CNN modules, RNN modules, attention modules, and a novel way of fusing different features”), in response to input of at least one of the one or plurality of pieces of whole information, at least one of the one or plurality of pieces of first part information, (Figure 2; Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “[…] Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted […]”; and Page 2, Feature fusion, 2nd Paragraph: “By extracting the pose keypoints from cropped pedestrian images, crossing behavior classifiers are built based on the human pose feature vectors. Improvement in prediction accuracy shows the effectiveness of using pose features”), the evaluation method comprising: (Figure 2; Page 2, Left Col., 3rd Paragraph: “The proposed method employs a novel neural network architecture for utilizing different spatio-temporal features with a hybrid fusion strategy”; Page 4, C. Model architecture, 1st Paragraph: “The overall architecture […] It consists of CNN modules, RNN modules, attention modules, and a novel way of fusing different features”; Page 2, Feature fusion, 2nd Paragraph: “By extracting the pose keypoints from cropped pedestrian images, crossing behavior classifiers are built based on the human pose feature vectors. Improvement in prediction accuracy shows the effectiveness of using pose features”; and Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “[…] Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted […]”); and (Figure 2; Page 2, Left Col., 3rd Paragraph: “The proposed method employs a novel neural network architecture for utilizing different spatio-temporal features with a hybrid fusion strategy”; Page 4, C. Model architecture, 1st Paragraph: “The overall architecture […] It consists of CNN modules, RNN modules, attention modules, and a novel way of fusing different features”; Page 3, III. A. Problem formulation, 1st – 2nd Paragraph: “[…] Given a sequence of observed video frames from the vehicle’s front view […] the goal is to design a model that can estimate the probability of the target pedestrian i’s action […] of crossing the road […] In the proposed model, explicit features such as pedestrian’s bounding box, pose keypoints, local context (cropped image around the pedestrian), and global context (semantic segmentation) are firstly extracted […]”; and Page 4, III. B. Input acquisition, Pedestrian pose keypoints: “Pedestrian pose keypoints represent the target pedestrian’s detailed motion, i.e., the posture at each frame while moving […] we utilize pre-trained OpenPose model to extract the pedestrian pose keypoints […] where p is a 36D vector of 2D coordinates that contain 18 pose joints […]”). Yang (2021) fails to teach one or plurality of pieces of second part information based on one or plurality of second parts of the movable object, the one or plurality of second parts being smaller than the one or plurality of first parts, and when a resolution of an image of the movable object detected from target image information is smaller than a threshold value, selecting the first model from the plurality of trained models; and when the resolution is equal to or greater than the threshold value, selecting the second model from the plurality of trained models. However, Rehder teaches one or plurality of pieces of second part information based on one or plurality of second parts of the movable object (Figure 1; Abstract: “to detect highly occluded pedestrians and estimate their head orientation. Detection is performed for pedestrian’s heads only”; Page 3, Left Col., 2nd Paragraph: “The second stage uses a part based model to which is trained only from detections of first stage. While complete pedestrians can vary greatly in pose and appearance, heads only vary little in appearance”; Page 3, Left Col., 6th Paragraph: “we focus on parts holding relevant information. We assign three part windows of identical size so that the information density per window is roughly the same and as large as possible”; and Page 3, Left Col., Last Paragraph: “[…] the part detectors are run on different positions within the bounding box of the first detector stage […]”), the one or plurality of second parts being smaller than the one or plurality of first parts (Page 2, III. A. Detection, 3rd Paragraph: “The first stage consists of a simple HOG/SVM classifier using a single head model at low resolution [...]”; Page 3, Left Col., 2nd Paragraph: “The second stage uses a part based model to which is trained only from detections of first stage. While complete pedestrians can vary greatly in pose and appearance, heads only vary little in appearance”; Page 3, Left Col., 6th Paragraph: “we focus on parts holding relevant information. We assign three part windows of identical size so that the information density per window is roughly the same and as large as possible”; Page 3, Left Col., Last Paragraph: “[…] the part detectors are run on different positions within the bounding box of the first detector stage […]”; and Page 2, Left Col., 3rd Paragraph: “[…] apply a root detector at low resolution and then detect parts of the object at double the root resolution while allowing for some deformation of the parts with respect to the root location”). Yang (2021) teaches evaluating a pedestrian's moving direction using whole body image information together with body part information obtained from human pose features, while Rehder teaches improving pedestrian analysis by performing hierarchical body part detection in which a larger detected body region is further analyzed using smaller constituent part regions through a second-stage part based model. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Rehder's hierarchical part-based representation into Yang's (2021) pedestrian intention prediction framework would have predictably enabled the use of progressively finer body-part information for movement evaluation while preserving the overall prediction architecture. The motivation for this combination of references would have been to further analyze a detected body region using a part based model and to detect parts of the object at double the root resolution, thereby utilizing progressively finer body part information to improve pedestrian movement evaluation. This motivation for the combination of Yang (2021) and Rehder is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Yang (2021) and Rehder fail to teach when a resolution of an image of the movable object detected from target image information is smaller than a threshold value, selecting the first model from the plurality of trained models; and when the resolution is equal to or greater than the threshold value, selecting the second model from the plurality of trained models. However, Yan teaches when a resolution of an image of the movable object detected from target image information is smaller than a threshold value(Figure 2; Abstract: “The serious performance decline with decreasing resolution is the major bottleneck for current pedestrian detection techniques. In this paper, we take pedestrian detection in different resolutions as different but related problems, and propose a Multi-Task model to jointly consider their commonness and differences”; Page 3035, Left Col., 2nd Paragraph: “[…] we consider the partition of two resolutions (low resolution: 30-80 pixels tall, and high resolution: taller than 80 pixels[…]”; and Page 3034, 3. Multi-Task Deformable Part Model, 1st Paragraph: “There are two intuitive strategies to handle the multiresolution detection. One is to combine samples from different resolutions to train a single detector (Fig. 2(a)), and another is to train independent detectors for different resolutions (Fig. 2(b))”). Therefore, it would have been obvious to one of ordinary skill in the art to combine Yang (2021), Rehder and Yan before the effective filing date of the claimed invention. The motivation for this combination of references would have been to utilize different spatio-temporal features with a hybrid fusion strategy (Yang 2021), to further analyze a detected body region using a part based model and detect parts of the object at double the root resolution (Rehder), and to retrieve higher-resolution measurements of the region so that the higher-resolution measurement can be used to extract a higher-accuracy object segment (Yan), thereby enabling progressively finer body-part information to be used for pedestrian movement evaluation according to the available image resolution. This motivation for the combination of Yang (2021), Rehder and Yan is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Yang (2021), Rehder and Yan fail to teach selecting the first model from the plurality of trained models; and selecting the second model from the plurality of trained models. However, Yang (2020) teaches selecting the first model (read as “lightweight sub-network”) from the plurality of trained models and selecting the second model (read as “higher resolution sub-networks”) from the plurality of trained models (Figure 1; Figure 2; Abstract: “In RANet, the input images are first routed to a lightweight sub-network that efficiently extracts low-resolution representations, and those samples with high prediction confidence will exit early from the network without being further processed. Meanwhile, high-resolution paths in the network maintain the capability to recognize the “hard” samples”; Page 3, 3.2. Overall Architecture, 1st Paragraph: “Figure 2 illustrates the overall architecture of the proposed RANet. It contains an Initial Layer and H subnetworks corresponding to different resolutions”; and Page 2, Left Col., 3rd Paragraph: “The “easy” samples are classified by the sub-network with the feature maps in the lowest spatial resolution. The sub-networks with higher resolution will be applied when the previous sub-network fails to achieve a given criterion”). Therefore, it would have been obvious to one of ordinary skill in the art to combine Yang (2021), Rehder, Yan and Yang (2020) before the effective filing date of the claimed invention. The motivation for this combination of references would have been to utilize different spatio-temporal features with a hybrid fusion strategy (Yang 2021), to further analyze detected body regions using a part based model and detect parts of the object at double the root resolution (Rehder), to retrieve higher-resolution measurements of the region so that the higher-resolution measurement can be used to extract a higher-accuracy object segment (Yan), and achieve a better trade-off between accuracy and computational cost by representations for easy samples while maintaining high-resolution processing capability for hard samples (Yang 2020), thereby enabling adaptive selection of an appropriate prediction model based on the available image information. This motivation for the combination of Yang (2021), Rehder, Yan and Yang (2020) is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Relevant Prior Art Directed to State of Art Shen et al (US 20200175326 A1) are relevant prior art not applied in the rejection(s) above. Shen discloses a method implemented by a system of one or more processors, the method comprising: obtaining an image from an image sensor of one or more image sensors positioned about a vehicle; determining a field of view for the image, the field of view being associated with a vanishing line; generating, from the image, a crop portion corresponding to the field of view, and a remaining portion, wherein the remaining portion of the image is downsampled; and outputting, via a convolutional neural network, information associated with detected objects depicted in the image, wherein detecting objects comprises performing a forward pass through the convolutional neural network of the crop portion and the remaining portion. Goel et al (US 12,100,224 B1) are relevant prior art not applied in the rejection(s) above. Goel discloses a system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising: receiving image data captured by an image sensor of a vehicle, the image data representing a portion of an environment in which the vehicle is operating; determining a first bounding box associated with a first pedestrian detected in the image data; determining, based at least in part on the first bounding box and the image data, first key points corresponding to physical features of the first pedestrian; inputting the image data into a machine-learned model; receiving an output from the machine-learned model, the output including second key points corresponding to physical features of a second pedestrian in the environment; determining, based at least in part on the second key points, a second bounding box associated with the second pedestrian, the second bounding box encompassing the second key points, the second bounding box indicating that the second pedestrian is in the environment; and controlling the vehicle based at least in part on the first bounding box and the second bounding box. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONGBONG NAH whose telephone number is (571) 272-1361. The examiner can normally be reached M - F: 9:00 AM - 5:30 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ONEAL MISTRY can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JONGBONG NAH/Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Oct 23, 2024
Application Filed
Jul 14, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705791
LOCALIZATION PROCESSING SERVICE
4y 2m to grant Granted Aug 11, 2026
Patent 12705905
ROAD BOUNDARY DETECTION BASED ON RADAR AND VISUAL INFORMATION
3y 9m to grant Granted Aug 11, 2026
Patent 12688721
DETECTING A CONDITION FOR A CULTURE DEVICE USING A MACHINE LEARNING MODEL
3y 11m to grant Granted Jul 21, 2026
Patent 12659419
AUGMENTED REALITY SELF-PORTRAITS
3y 11m to grant Granted Jun 16, 2026
Patent 12645937
SYSTEM, METHOD, AND COMPUTER PROGRAM FOR ITERATIVE CONTENT ADAPTIVE ONLINE TRAINING IN NEURAL IMAGE COMPRESSION
3y 8m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
76%
Grant Probability
93%
With Interview (+16.3%)
2y 10m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 118 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month