Prosecution Insights
Last updated: October 02, 2026
Application No. 19/014,083

IMAGE SEQUENCE DETECTION METHOD AND APPARATUS, MEDIUM, DEVICE, AND PROGRAM PRODUCT

Non-Final OA §103
Filed
Jan 08, 2025
Priority
Jan 17, 2023 — CN 202310090615.7 +1 more
Examiner
MANGIALASCHI, TRACY
Art Unit
Tech Center
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
1 (Non-Final)
76%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
455 granted / 603 resolved
+15.5% vs TC avg
Strong +27% interview lift
Without
With
+27.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
13 currently pending
Career history
614
Total Applications
across all art units

Statute-Specific Performance

§101
8.6%
-31.4% vs TC avg
§103
55.9%
+15.9% vs TC avg
§102
14.2%
-25.8% vs TC avg
§112
15.9%
-24.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 603 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of the Claims Claims 1-20, as originally filed, are currently pending and have been considered below. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xu et al., Chinese Publication No. CN 112827168 A, hereinafter, “Xu”, and further in view of Zhao et al., Chinese Publication No. CN 113627334 A, hereinafter, “Zhao”. As per claim 1, Xu discloses an image sequence detection method performed by a computer device, the method comprising: obtaining an initial image sequence, the initial image sequence comprising a target object (Xu, ¶n0005, This application provides a method, apparatus, and storage medium for target tracking. It can determine multiple candidate boxes in the game image to be identified based on the difference between the game image to be identified and historical game images; Xu, ¶n0007, Acquire the game image to be identified and historical game images, wherein the game image to be identified is the next frame image adjacent to the historical game image; Xu, ¶n0012, Based on the target object to be tracked, obtain the game image to be identified and the historical game image about the target object; wherein, the game image to be identified is the next frame image adjacent to the historical game image; Xu, ¶n0099, 901. The server obtains the game image to be identified and historical game images); inputting the initial image sequence to a pretrained target detection model, to obtain a detection result sequence (Xu, ¶n0022, the target tracking device also includes a processing unit for inputting the spatial features corresponding to each candidate box into a classifier; Xu, ¶n0058, This application relates to the field of computer vision ... computer vision ... to build artificial intelligence systems capable of extracting information from images; Xu, ¶n0079, Once the training corresponding to step S4 is completed, new classifier parameters are obtained, thus yielding the target tracking model; Xu, ¶n0080-n0082, Please refer to Figure 5 ... As shown in the figure, an embodiment of the target tracking method in this application includes: 501. The server obtains the game image to be identified and historical game images. The game image to be identified is the next frame image adjacent to the historical game image. When the server receives an instruction to track a target object, it needs to obtain the game screen corresponding to the current game video frame and input it as the game image to be identified into the target tracking model. At the same time, it is also necessary to acquire historical game video frames. Since the game video screen is continuous and the movement trajectory of the target object is also continuous, historical game video frames are a prerequisite for the server to identify the position of the target object in the current game image to be identified. The position of the target object in the current game screen can be predicted by the target object in the historical game video frames; Xu, ¶n0105, 904. The server uses the first convolution operator to extract convolution features from the game image to be identified, thus obtaining the convolution features of the game image to be identified; Xu, ¶n0113, 908. The server obtains multiple candidate features from the color difference map of each candidate box and inputs them into the classifier; Xu, ¶n0118, through multiple convolution operations, the image feature differences of the game image to be identified are amplified. Candidate features of each candidate box are extracted based on the color difference map. The probability that the image corresponding to each candidate box is the target object is determined by comparing the candidate features. Finally, the target candidate box is determined based on the probability score of each candidate box, and the location of the target candidate box is the location of the target object); and determining a detection result corresponding to the initial image sequence based on the detection result sequence, the detection result indicating a category of the target object (Xu, ¶n0015, The location information of the target object is determined based on the location information corresponding to the target bounding box; Xu, ¶n0060, server needs to identify and track the target object in the game video screen, determine the location of the target object, and finally automatically complete the shooting of the target object based on the location; Xu, ¶n0100, The game image to be identified is the next frame image adjacent to the historical game image. When the server receives a tracking instruction for a target object, it needs to obtain the game screen corresponding to the current game video frame and input it as the game image to be identified into the target tracking model. At the same time, the server also needs to acquire historical game video frames. Since the game video is continuous and the movement trajectory of the target object is also continuous, historical game videoframes are a prerequisite for the server to identify the target being tracked in the current game image. The location of the target object in the current game image can be predicted by the images of historical game video frames; Xu, ¶n0104, The movement of the target object will cause a frame difference between the historical game video image and the game image to be identified. When there is a significant change between the pixels of consecutive frames, the frame difference between consecutive frames can be used to determine the change area; Xu, ¶n0115, 909. The server obtains the probability score corresponding to each candidate box through the classifier, and determines the target box based on the probability score of each candidate box). Xu does not explicitly disclose the following limitation as further recited however Zhao discloses determining a detection result corresponding to each image in the initial image sequence based on the detection result sequence, the detection result indicating a category of the target object (Zhao, ¶n0033, Step S202: Detect key points of the target object in multiple consecutive frames of images to obtain a key point sequence of the target object, wherein the key point sequence includes the key points of the target object in the multiple consecutive frames of images; Zhao, ¶n0035, Step S204: Analyze the key point sequence to obtain the static features and dynamic features of the target object, wherein the static features represent the positional relationship of different key points of the target object in the same frame image, and the dynamic features represent the positional relationship of the same key points of the target object in different frame images; Zhao, ¶n0037, Step S206: Based on the static features and the dynamic features, identify the behavior of the target object; Zhao, ¶n0008, obtaining the behavior recognition result of the target object through the target static features and the target dynamic features includes: inputting the target static features and the target dynamic features into the fully connected layer of the behavior recognition neural network model; analyzing the target static features and the target dynamic features through the fully connected layer to obtain the behavior category of the target object, wherein the behavior recognition result of the target object includes the behavior category of the target object). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhao with Xu because they are in the same field of endeavor. One skilled in the art would have been motivated to include the use of a sequence of consecutive frames to determine a detection result as taught by Zhao in the system of Xu in order to provide an alternate means to recognize targets in adjacent and consecutive frames (Zhao, ¶0001). As per claim 2, Xu and Zhao disclose the method according to claim 1, wherein the target detection model is obtained by inputting sample training data to an initial detection model to be trained for training, the sample training data comprising one set of sample feature information respectively extracted from one set of sample images (Xu, ¶n0030, the acquisition unit is also used to acquire positive samples in the game image to be trained; the positive samples are candidate boxes where the tracked object is located in the game mage to be trained; and a first loss value is determined based on the spatial feature set corresponding to the positive samples and the spatial feature set of positive samples in historical game images; Xu, ¶n0075-n0079, The training process may include the following steps: In step S1, a game recording sample is obtained … which includes game characters that need to be automatically identified and tracked … In step S2, the game video frame obtained from the game recording sample in step S1 is used as the input to the target tracking model to be trained. The classifier in the target tracking model outputs a prediction sample, which includes the feature set of multiple prediction boxes in the game frame … In step S3, the predicted samples obtained in step S2 are used as the input of the temporal focusing model. The temporal focusing model also needs to obtain the predicted samples corresponding to the historical game frames before the current game frame … In step S4, the loss value obtained in step S2 is used to train the operational parameters of the classifier. Once the training corresponding to step S4 is completed, new classifier parameters are obtained, thus yielding the target tracking model), and a classification result and an evaluation parameter that correspond to each piece of sample feature information, the classification result indicating a classification category for one piece of sample feature information, the evaluation parameter indicating a reward corresponding to the classification category, the one set of sample images being one set of continuous frames of images, and for an ith frame of sample image in the one set of continuous frames of images, sample feature information of the ith frame of sample image being jointly determined based on the ith frame of sample image and an (i−1)th frame of sample image (Xu, ¶n0022, the target tracking device also includes a processing unit for inputting the spatial features corresponding to each candidate box into a classifier, and outputting the probability score corresponding to each candidate box through the classifier; Xu, ¶n0031, The processing unit is also used to update the operational parameters of the classifier to be trained based on the first loss value; when the classifier update condition is met, the classifier is obtained based on the updated operational parameters of the classifier to be trained; Xu, ¶n0073, The classifier is used to perform calculations on image features to determine whether the image in the candidate box is the target object. For example, the classifier can output the probability that the image in each candidate box is the target object, and the target candidate box containing the target object can be determined based on the probability score of each candidate box output by the classifier ... a classifier needs to analyze whether an image in a candidate box is a target object based on the weights of different features. Therefore, it needs to continuously adjust the classifier's operational parameters to adjust the proportion of each image feature in target recognition; Xu, ¶n0074, The temporal focus model uses the features of candidate boxes in the current game image to be identified and the features of candidate boxes in historical game images to determine the image features that are more likely to predict the target object in the next frame of the game image. Based on these image features, the weights of each feature in the classifier are adjusted to improve the classifier's recognition accuracy for the next frame of the game image ... the temporal focus model can continuously learn from the recognition results of historical game images to be identified, obtain the image features most relevant to the recognition of the target object in the next frame of the game image, and continuously adjust the weight of the most relevant image features to improve the accuracy of the classifier in recognizing and tracking the next frame of the game image). As per claim 3, Xu and Zhao disclose the method according to claim 2, further comprising: performing a first feature extraction operation on the one set of sample images, to obtain a first set of feature information, feature information in the first set of feature information being in one-to-one correspondence to sample images in the one set of sample images (Xu, ¶n0009, Obtain the spatial features corresponding to each of at least two candidate boxes, wherein the spatial features include the position information of the candidate box in the game image to be identified and the color difference map corresponding to the candidate box); performing a second feature extraction operation on the one set of sample images, to obtain a second set of feature information, first feature information in the second set of feature information being in one-to-one correspondence to the sample images in the one set of sample images, and the first feature information of the ith frame of sample image indicating a correlation feature between the ith frame of sample image and the (i−1)th frame of sample image (Xu, ¶n0027, the acquisition unit is also used to acquire multiple candidate features from the color difference map and determine the weight value of each candidate feature; Xu, ¶n0074, The temporal focus model uses the features of candidate boxes in the current game image to be identified and the features of candidate boxes in historical game images to determine the image features that are more likely to predict the target object in the next frame of the game image; Xu, ¶n0095, the pooling features corresponding to the game image to be identified need to be input into a special target convolutional layer for feature extraction again; Xu, ¶n0096, the server determines the candidate box based on the frame difference between the previous and next game screens); and fusing the first set of feature information and the second set of feature information respectively, to obtain the one set of sample feature information (Xu, ¶n0025, the acquisition unit is specifically used to extract convolutional features from the game image to be identified using a first convolution operator to obtain the convolutional features of the game image to be identified; to perform pooling processing on the convolutional features of the game image to be identified to obtain the pooling features of the game image to be identified; to determine the target pooling feature in the general pooling features based on the position information of at least two candidate boxes; and to determine the color difference map corresponding to each candidate box in the at least two candidate boxes based on the target pooling feature; Xu, ¶n0096, the second convolutional features obtained from the special target convolutional layer can be concatenated using a fully connected layer, and then a specific low-frequency convolutional operator(third convolutional layer) can be used for convolutional reconstruction to reconstruct a visualized image, further expanding the spatial differences. Finally, a color difference map is obtained, as shown in Figure 8; Xu, ¶n0113, 908. The server obtains multiple candidate features from the color difference map of each candidate box and inputs them into the classifier). As per claim 4, Xu and Zhao disclose the method according to claim 3, wherein for the ith frame of sample image, the performing a second feature extraction operation on the one set of sample images, to obtain a second set of feature information comprises: performing an image recognition operation on the ith frame of sample image, to determine a first set of key points, the first set of key points being configured for describing a location of the target object in the ith frame of sample image (Zhao, ¶n0006, detecting key points of a target object in multiple consecutive frames of images to obtain a key point sequence of the target object, wherein the key point sequence includes key points of the target object in the multiple consecutive frames of image; analyzing the key point sequence to obtain static features and dynamic features of the target object, wherein the static features represent the positional relationship of different key points of the target object in the same frame of images, and the dynamic features represent the positional relationship of the same key points of the target object in different frames of images); performing an image recognition operation on the (i−1)th frame of sample image, to determine a second set of key points, the second set of key points being configured for describing a location of the target object in the (i−1)th frame of sample image (Zhao, ¶n0006, detecting key points of a target object in multiple consecutive frames of images to obtain a key point sequence of the target object, wherein the key point sequence includes key points of the target object in the multiple consecutive frames of image; analyzing the key point sequence to obtain static features and dynamic features of the target object, wherein the static features represent the positional relationship of different key points of the target object in the same frame of images, and the dynamic features represent the positional relationship of the same key points of the target object in different frames of images); and determining the first feature information of the ith frame of sample image based on the first set of key points and the second set of key points, the first feature information of the ith frame of sample image being configured for indicating a distance between corresponding key points in the first set of key points and the second set of key points (Zhao, ¶n0006, detecting key points of a target object in multiple consecutive frames of images to obtain a key point sequence of the target object, wherein the key point sequence includes key points of the target object in the multiple consecutive frames of image; analyzing the key point sequence to obtain static features and dynamic features of the target object, wherein the static features represent the positional relationship of different key points of the target object in the same frame of images, and the dynamic features represent the positional relationship of the same key points of the target object in different frames of images; Zhao, ¶n0010, the key point sequence is analyzed to obtain the static features of the target object, including: determining the distance code between any two key points among multiple key points in each frame image, and obtaining the distance-coded features in the static features; Zhao, ¶n0069, the dynamic feature can be the trajectory coding feature between the same key point in different frame images or the trajectory direction coding feature between the same key point in different frame images ... The trajectory encoding feature is used to represent the motion trajectory information of the same key point in different frame images. It can be represented by the difference of the coordinate points of the same key point in multiple selected image frames. For example, the trajectory information of the key point can be represented by the difference of the coordinate points of the same key point in three selected image frames. The trajectory direction encoding is used to represent the motion trajectory direction information of the same key point in different frame images. It can be represented by the difference of the coordinate points of the same key point in multiple selected image frames. For example, the trajectory direction information of the key point can be represented by the difference of the coordinate points of the same key point in three selected image frames). As pr claim 5, Xu and Zhao disclose the method according to claim 2, wherein for the ith frame of sample image, the method further comprises: inputting the sample feature information of the ith frame of sample image to the initial detection model, to obtain a classification result for the ith frame of sample image, the initial detection model being a detection model obtained by performing an initialization operation in advance (Xu, ¶n0022, the target tracking device also includes a processing unit for inputting the spatial features corresponding to each candidate box into a classifier, and outputting the probability score corresponding to each candidate box through the classifier; Xu, ¶n0073, The classifier is used to perform calculations on image features to determine whether the image in the candidate box is the target object. For example, the classifier can output the probability that the image in each candidate box is the target object, and the target candidate box containing the target object can be determined based on the probability score of each candidate box output by the classifier ... a classifier needs to analyze whether an image in a candidate box is a target object based on the weights of different features); and determining an evaluation parameter for the ith frame of sample image based on the classification result for the ith frame of sample image and a classification category corresponding to the sample feature information of the ith frame of sample image, the evaluation parameter for the ith frame of sample image being configured for indicating whether the classification result for the ith frame of sample image is the same as the corresponding classification category (Xu, ¶n0031, The processing unit is also used to update the operational parameters of the classifier to be trained based on the first loss value; when the classifier update condition is met, the classifier is obtained based on the updated operational parameters of the classifier to be trained; Xu, ¶n0074, The temporal focus model uses the features of candidate boxes in the current game image to be identified and the features of candidate boxes in historical game images to determine the image features that are more likely to predict the target object in the next frame of the game image. Based on these image features, the weights of each feature in the classifier are adjusted to improve the classifier's recognition accuracy for the next frame of the game image). As per claim 6, Xu and Zhao disclose the method according to claim 5, wherein the inputting the sample feature information of the ith frame of sample image to the initial detection model, to obtain a classification result for the ith frame of sample image comprises: inputting each category in a preset category set to a first preset function in sequence with the sample feature information of the ith frame of sample image, to obtain a first set of function values, the first set of function values comprising a function value for each category relative to the sample feature information of the ith frame of sample image (Xu, ¶n0073, The classifier is used to perform calculations on image features to determine whether the image in the candidate box is the target object. For example, the classifier can output the probability that the image in each candidate box is the target object, and the target candidate box containing the target object can be determined based on the probability score of each candidate box output by the classifier; Xu, ¶n0077, The classifier in the target tracking model outputs a prediction sample, which includes the feature set of multiple prediction boxes in the game frame and the probability score of the image corresponding to each prediction box being the game character to be tracked; Xu, ¶n0088, The classifier classifies each candidate box based on the spatial features corresponding to each candidate box. That is, it classifies and scores each candidate box based on multiple image features, and obtains the probability score of the image corresponding to each candidate box being the target object. Then, it outputs the probability score of each candidate box, and then uses the probability score to finally determine the recognition result); and determining, as the classification result for the ith frame of sample image, a candidate category corresponding to a first function value in the first set of function values that is the largest (Xu, ¶n0090, After the classifier outputs the probability score of each candidate box, it needs to determine whether the probability score of each candidate box exceeds a preset threshold. If the probability scores of all candidate boxes do not exceed the preset threshold, it means that the target object has not been identified, and the target tracking model has failed to track the target object. If the probability scores of multiple candidate boxes all exceed the preset threshold, then the candidate box with the highest average probability needs to be selected as the final recognition result. That is, the image corresponding to the candidate box is recognized as the target object). As per claim 7, Xu and Zhao disclose the method according to claim 6, wherein before the inputting each category in a preset category set to a first preset function in sequence with the sample feature information of the ith frame of sample image, to obtain a first set of function values, the method further comprises: calculating a target random number in a preset manner (Zhao, ¶n0034, the consecutive multi-frame images can be any consecutive images in the video stream captured by the monitoring equipment, or any consecutive images in the video file stored in the computer or mobile phone ... The target object can be people, animals, etc. in a series of images; Zhao, ¶n0038, the behavior of the target object can be identified by the static and dynamic features of the target object. The result of behavior identification can be the type of human behavior, which can be preset); determining whether the target random number satisfies a preset probability condition (Zhao, ¶n0038, the behavior of the target object can be identified by the static and dynamic features of the target object. The result of behavior identification can be the type of human behavior, which can be preset ... The specific type can be determined according to the actual situation); and selecting a candidate category randomly from the preset category set as the classification result for the ith frame of sample image when the target random number satisfies the preset probability condition; or performing the operation of inputting each category in a preset category set to a first preset function in sequence with the sample feature information of the ith frame of sample image, to obtain a first set of function values, when the target random number does not satisfy the preset probability condition (Zhao, ¶n0034, the consecutive multi-frame images can be any consecutive images in the video stream captured by the monitoring equipment, or any consecutive images in the video file stored in the computer or mobile phone ... The target object can be people, animals, etc. in a series of images; Zhao, ¶n0006, a method for recognizing the behavior of an object is provided, comprising: detecting key points of a target object in multiple consecutive frames of images to obtain a key point sequence of the target object, wherein the key point sequence includes keypoints of the target object in the multiple consecutive frames of images; analyzing the key point sequence to obtain static features and dynamic features of the target object, wherein the static features represent the positional relationship of different key points of the target object in the same frame of images, and the dynamic features represent the positional relationship of the same key points of the target object in different frames of images; and recognizing the behavior of the target object based on the static features and the dynamic features; Zhao, ¶n0008, obtaining the behavior recognition result of the target object through the target static features and the target dynamic features includes: inputting the target static features and the target dynamic features into the fully connected layer of the behavior recognition neural network model; analyzing the target static features and the target dynamic features through the fully connected layer to obtain the behavior category of the target object, wherein the behavior recognition result of the target object includes the behavior category of the target object). As per claim 8, Xu and Zhao disclose the method according to claim 6, wherein the determining, as the classification result for the ith frame of sample image, a candidate category corresponding to a first function value in the first set of function values that is the largest comprises: determining the evaluation parameter for the ith frame of sample image as a first evaluation parameter when the candidate category corresponding to the first function value is the same as the classification category for the ith frame of sample image, the first evaluation parameter indicating correct detection of the initial detection model (Xu, ¶n0090, After the classifier outputs the probability score of each candidate box, it needs to determine whether the probability score of each candidate box exceeds a preset threshold. If the probability scores of all candidate boxes do not exceed the preset threshold, it means that the target object has not been identified, and the target tracking model has failed to track the target object. If the probability scores of multiple candidate boxes all exceed the preset threshold, then the candidate box with the highest average probability needs to be selected as the final recognition result. That is, the image corresponding to the candidate box is recognized as the target object); or determining the evaluation parameter for the ith frame of sample image as a second evaluation parameter when the candidate category corresponding to the first function value is different from the classification category for the ith frame of sample image, the second evaluation parameter indicating incorrect detection of the initial detection model (Xu, ¶n0090, After the classifier outputs the probability score of each candidate box, it needs to determine whether the probability score of each candidate box exceeds a preset threshold. If the probability scores of all candidate boxes do not exceed the preset threshold, it means that the target object has not been identified, and the target tracking model has failed to track the target object. If the probability scores of multiple candidate boxes all exceed the preset threshold, then the candidate box with the highest average probability needs to be selected as the final recognition result. That is, the image corresponding to the candidate box is recognized as the target object). As per claim 9, Xu and Zhao disclose the method according to claim 2, wherein the target detection model is obtained by: selecting a plurality of sample data structures from the sample training data, a jth sample data structure in the plurality of sample data structures comprising a jth piece of sample feature information, a jth classification result corresponding to the jth piece of sample feature information, a jth evaluation parameter, and a (j+1)th piece of sample feature information, and j being a positive integer (Xu, the acquisition unit is also used to acquire positive samples in the game image to be trained; the positive samples are candidate boxes where the tracked object is located in the game mage to be trained; and a first loss value is determined based on the spatial feature set corresponding to the positive samples and the spatial feature set of positive samples in historical game images); and training the initial detection model to be trained based on the plurality of sample data structures, to obtain the target detection model, the initial detection model being determined as the target detection model when a round count of training the initial detection model reaches a preset round count threshold, or a model parameter of the initial detection model being adjusted based on a predetermined loss function when a round count of training the initial detection model does not reach the preset round count threshold, and an input for a training process of each round being one of the plurality of sample data structures (Xu, ¶n0031, The processing unit is also used to update the operational parameters of the classifier to be trained based on the first loss value; when the classifier update condition is met, the classifier is obtained based on the updated operational parameters of the classifier to be trained). As per claim 10, Xu and Zhao disclose the method according to claim 9, wherein for the jth sample data structure, the training the initial detection model to be trained based on the plurality of sample data structures, to obtain the target detection model comprises: determining whether a sample image corresponding to the (j+1)th piece of sample feature information is a last sample image in the sample image sequence (Xu, ¶n0030, the acquisition unit is also used to acquire positive samples in the game image to be trained; the positive samples are candidate boxes where the tracked object is located in the game mage to be trained; and a first loss value is determined based on the spatial feature set corresponding to the positive samples and the spatial feature set of positive samples in historical game images; Xu, ¶n0031, The processing unit is also used to update the operational parameters of the classifier to be trained based on the first loss value; when the classifier update condition is met, the classifier is obtained based on the updated operational parameters of the classifier to be trained; Xu, ¶n0032, an acquisition unit acquires positive and negative samples from the game image to be trained; positive samples are candidate boxes where the tracked object is located in the game image to be trained, and negative samples are candidate boxes that do not include the tracked image; a second loss value is determined based on the spatial feature set corresponding to the positive samples, the spatial feature set corresponding to the negative samples, and the spatial feature set of the historical game images); determining, based on the jth evaluation parameter when the sample image corresponding to the (j+1)th piece of sample feature information is the last sample image, a preset parameter outputted by the initial detection model; or determining the preset parameter based on the jth evaluation parameter and an output of a reference detection model when the sample image corresponding to the (j+1)th piece of sample feature information is not the last sample image, the reference detection model being a detection model obtained by performing an initialization operation in advance, and a parameter of the reference detection model being different from a parameter of the initial detection model (Xu, ¶n0033, The determination unit is also used to update the operational parameters of the classifier to be trained based on the second loss value; when the classifier update condition is met, the classifier is obtained based on the updated operational parameters of the classifier to be trained); determining a value of the loss function based on the preset parameter and an output of the initial detection model, and updating the parameter of the initial detection model based on the value of the loss function (Xu, ¶n0033, The determination unit is also used to update the operational parameters of the classifier to be trained based on the second loss value; when the classifier update condition is met, the classifier is obtained based on the updated operational parameters of the classifier to be trained); and determining the initial detection model as the target detection model when a round count of performing the above operations reaches the preset round count threshold (Xu, ¶n0033, The determination unit is also used to update the operational parameters of the classifier to be trained based on the second loss value; when the classifier update condition is met, the classifier is obtained based on the updated operational parameters of the classifier to be trained; Xu, ¶n0079, In step S4, the loss value obtained in step S2 is used to train the operational parameters of the classifier. Once the training corresponding to step S4 is completed, new classifier parameters are obtained, thus yielding the target tracking model). As per claim 11, Xu disclose an electronic device, comprising a memory and a processor, the memory having computer programs stored therein, and the computer programs, when executed by the processor, causing the electronic device to perform an image sequence detection method including: obtaining an initial image sequence, the initial image sequence comprising a target object (Xu, ¶n0005, This application provides a method, apparatus, and storage medium for target tracking. It can determine multiple candidate boxes in the game image to be identified based on the difference between the game image to be identified and historical game images; Xu, ¶n0007, Acquire the game image to be identified and historical game images, wherein the game image to be identified is the next frame image adjacent to the historical game image; Xu, ¶n0012, Based on the target object to be tracked, obtain the game image to be identified and the historical game image about the target object; wherein, the game image to be identified is the next frame image adjacent to the historical game image; Xu, ¶n0099, 901. The server obtains the game image to be identified and historical game images); inputting the initial image sequence to a pretrained target detection model, to obtain a detection result sequence (Xu, ¶n0022, the target tracking device also includes a processing unit for inputting the spatial features corresponding to each candidate box into a classifier; Xu, ¶n0058, This application relates to the field of computer vision ... computer vision ... to build artificial intelligence systems capable of extracting information from images; Xu, ¶n0079, Once the training corresponding to step S4 is completed, new classifier parameters are obtained, thus yielding the target tracking model; Xu, ¶n0080-n0082, Please refer to Figure 5 ... As shown in the figure, an embodiment of the target tracking method in this application includes: 501. The server obtains the game image to be identified and historical game images. The game image to be identified is the next frame image adjacent to the historical game image. When the server receives an instruction to track a target object, it needs to obtain the game screen corresponding to the current game video frame and input it as the game image to be identified into the target tracking model. At the same time, it is also necessary to acquire historical game video frames. Since the game video screen is continuous and the movement trajectory of the target object is also continuous, historical game video frames are a prerequisite for the server to identify the position of the target object in the current game image to be identified. The position of the target object in the current game screen can be predicted by the target object in the historical game video frames; Xu, ¶n0105, 904. The server uses the first convolution operator to extract convolution features from the game image to be identified, thus obtaining the convolution features of the game image to be identified; Xu, ¶n0113, 908. The server obtains multiple candidate features from the color difference map of each candidate box and inputs them into the classifier; Xu, ¶n0118, through multiple convolution operations, the image feature differences of the game image to be identified are amplified. Candidate features of each candidate box are extracted based on the color difference map. The probability that the image corresponding to each candidate box is the target object is determined by comparing the candidate features. Finally, the target candidate box is determined based on the probability score of each candidate box, and the location of the target candidate box is the location of the target object); and determining a detection result corresponding to the initial image sequence based on the detection result sequence, the detection result indicating a category of the target object(Xu, ¶n0015, The location information of the target object is determined based on the location information corresponding to the target bounding box; Xu, ¶n0060, server needs to identify and track the target object in the game video screen, determine the location of the target object, and finally automatically complete the shooting of the target object based on the location; Xu, ¶n0100, The game image to be identified is the next frame image adjacent to the historical game image. When the server receives a tracking instruction for a target object, it needs to obtain the game screen corresponding to the current game video frame and input it as the game image to be identified into the target tracking model. At the same time, the server also needs to acquire historical game video frames. Since the game video is continuous and the movement trajectory of the target object is also continuous, historical game videoframes are a prerequisite for the server to identify the target being tracked in the current game image. The location of the target object in the current game image can be predicted by the images of historical game video frames; Xu, ¶n0104, The movement of the target object will cause a frame difference between the historical game video image and the game image to be identified. When there is a significant change between the pixels of consecutive frames, the frame difference between consecutive frames can be used to determine the change area; Xu, ¶n0115, 909. The server obtains the probability score corresponding to each candidate box through the classifier, and determines the target box based on the probability score of each candidate box). Xu does not explicitly disclose the following limitation as further recited however Zhao discloses determining a detection result corresponding to each image in the initial image sequence based on the detection result sequence, the detection result indicating a category of the target object (Zhao, ¶n0033, Step S202: Detect key points of the target object in multiple consecutive frames of images to obtain a key point sequence of the target object, wherein the key point sequence includes the key points of the target object in the multiple consecutive frames of images; Zhao, ¶n0035, Step S204: Analyze the key point sequence to obtain the static features and dynamic features of the target object, wherein the static features represent the positional relationship of different key points of the target object in the same frame image, and the dynamic features represent the positional relationship of the same key points of the target object in different frame images; Zhao, ¶n0037, Step S206: Based on the static features and the dynamic features, identify the behavior of the target object; Zhao, ¶n0008, obtaining the behavior recognition result of the target object through the target static features and the target dynamic features includes: inputting the target static features and the target dynamic features into the fully connected layer of the behavior recognition neural network model; analyzing the target static features and the target dynamic features through the fully connected layer to obtain the behavior category of the target object, wherein the behavior recognition result of the target object includes the behavior category of the target object). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhao with Xu because they are in the same field of endeavor. One skilled in the art would have been motivated to include the use of a sequence of consecutive frames to determine a detection result as taught by Zhao in the system of Xu in order to provide an alternate means to recognize targets in adjacent and consecutive frames (Zhao, ¶0001). As per claim 19, Xu discloses a non-transitory computer-readable storage medium, comprising computer programs therein, wherein the computer programs, when executed by a processor of an electronic device, cause the electronic device to perform an image sequence detection method including: obtaining an initial image sequence, the initial image sequence comprising a target object (Xu, ¶n0005, This application provides a method, apparatus, and storage medium for target tracking. It can determine multiple candidate boxes in the game image to be identified based on the difference between the game image to be identified and historical game images; Xu, ¶n0007, Acquire the game image to be identified and historical game images, wherein the game image to be identified is the next frame image adjacent to the historical game image; Xu, ¶n0012, Based on the target object to be tracked, obtain the game image to be identified and the historical game image about the target object; wherein, the game image to be identified is the next frame image adjacent to the historical game image; Xu, ¶n0099, 901. The server obtains the game image to be identified and historical game images); inputting the initial image sequence to a pretrained target detection model, to obtain a detection result sequence (Xu, ¶n0022, the target tracking device also includes a processing unit for inputting the spatial features corresponding to each candidate box into a classifier; Xu, ¶n0058, This application relates to the field of computer vision ... computer vision ... to build artificial intelligence systems capable of extracting information from images; Xu, ¶n0079, Once the training corresponding to step S4 is completed, new classifier parameters are obtained, thus yielding the target tracking model; Xu, ¶n0080-n0082, Please refer to Figure 5 ... As shown in the figure, an embodiment of the target tracking method in this application includes: 501. The server obtains the game image to be identified and historical game images. The game image to be identified is the next frame image adjacent to the historical game image. When the server receives an instruction to track a target object, it needs to obtain the game screen corresponding to the current game video frame and input it as the game image to be identified into the target tracking model. At the same time, it is also necessary to acquire historical game video frames. Since the game video screen is continuous and the movement trajectory of the target object is also continuous, historical game video frames are a prerequisite for the server to identify the position of the target object in the current game image to be identified. The position of the target object in the current game screen can be predicted by the target object in the historical game video frames; Xu, ¶n0105, 904. The server uses the first convolution operator to extract convolution features from the game image to be identified, thus obtaining the convolution features of the game image to be identified; Xu, ¶n0113, 908. The server obtains multiple candidate features from the color difference map of each candidate box and inputs them into the classifier; Xu, ¶n0118, through multiple convolution operations, the image feature differences of the game image to be identified are amplified. Candidate features of each candidate box are extracted based on the color difference map. The probability that the image corresponding to each candidate box is the target object is determined by comparing the candidate features. Finally, the target candidate box is determined based on the probability score of each candidate box, and the location of the target candidate box is the location of the target object); and determining a detection result corresponding to the initial image sequence based on the detection result sequence, the detection result indicating a category of the target object (Xu, ¶n0015, The location information of the target object is determined based on the location information corresponding to the target bounding box; Xu, ¶n0060, server needs to identify and track the target object in the game video screen, determine the location of the target object, and finally automatically complete the shooting of the target object based on the location; Xu, ¶n0100, The game image to be identified is the next frame image adjacent to the historical game image. When the server receives a tracking instruction for a target object, it needs to obtain the game screen corresponding to the current game video frame and input it as the game image to be identified into the target tracking model. At the same time, the server also needs to acquire historical game video frames. Since the game video is continuous and the movement trajectory of the target object is also continuous, historical game videoframes are a prerequisite for the server to identify the target being tracked in the current game image. The location of the target object in the current game image can be predicted by the images of historical game video frames; Xu, ¶n0104, The movement of the target object will cause a frame difference between the historical game video image and the game image to be identified. When there is a significant change between the pixels of consecutive frames, the frame difference between consecutive frames can be used to determine the change area; Xu, ¶n0115, 909. The server obtains the probability score corresponding to each candidate box through the classifier, and determines the target box based on the probability score of each candidate box). Xu does not explicitly disclose the following limitation as further recited however Zhao discloses determining a detection result corresponding to each image in the initial image sequence based on the detection result sequence, the detection result indicating a category of the target object (Zhao, ¶n0033, Step S202: Detect key points of the target object in multiple consecutive frames of images to obtain a key point sequence of the target object, wherein the key point sequence includes the key points of the target object in the multiple consecutive frames of images; Zhao, ¶n0035, Step S204: Analyze the key point sequence to obtain the static features and dynamic features of the target object, wherein the static features represent the positional relationship of different key points of the target object in the same frame image, and the dynamic features represent the positional relationship of the same key points of the target object in different frame images; Zhao, ¶n0037, Step S206: Based on the static features and the dynamic features, identify the behavior of the target object; Zhao, ¶n0008, obtaining the behavior recognition result of the target object through the target static features and the target dynamic features includes: inputting the target static features and the target dynamic features into the fully connected layer of the behavior recognition neural network model; analyzing the target static features and the target dynamic features through the fully connected layer to obtain the behavior category of the target object, wherein the behavior recognition result of the target object includes the behavior category of the target object). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhao with Xu because they are in the same field of endeavor. One skilled in the art would have been motivated to include the use of a sequence of consecutive frames to determine a detection result as taught by Zhao in the system of Xu in order to provide an alternate means to recognize targets in adjacent and consecutive frames (Zhao, ¶0001). Regarding claim(s) 12 and 20: A corresponding reasoning as given earlier (see rejection of claim(s) 2) applies, mutatis mutandis, to the subject-matter of claim(s) 12 and 20, and therefore is/are also considered rejected under the grounds given in the rejection of claim(s) 2. Regarding claim(s) 13-16: A corresponding reasoning as given earlier (see rejection of claim(s) 3-6) applies, mutatis mutandis, to the subject-matter of claim(s) 13-16, and therefore is/are also considered rejected under the grounds given in the rejection of claim(s) 3-6. Regarding claim(s) 17 and 18: A corresponding reasoning as given earlier (see rejection of claim(s) 9 and 10) applies, mutatis mutandis, to the subject-matter of claim(s) 17 and 18, and therefore is/are also considered rejected under the grounds given in the rejection of claim(s) 9 and 10. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to TRACY MANGIALASCHI whose telephone number is (571)270-5189. The examiner can normally be reached M-F, 9:30AM TO 6:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TRACY MANGIALASCHI/Primary Examiner, Art Unit 2668
Read full office action

Prosecution Timeline

Jan 08, 2025
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737917
COMBINING DOMAIN KNOWLEDGE AND FOUNDATION MODELS FOR ONE-SHOT MEDICAL IMAGE FEATURE LOCALIZATION
2y 6m to grant Granted Sep 15, 2026
Patent 12733729
METHOD AND SYSTEM FOR DETERMINING COSMETIC SKIN ATTRIBUTES BASED ON DISORDER VALUE
2y 3m to grant Granted Sep 15, 2026
Patent 12731410
METHODS AND APPARATUS TO TRACK AND CLASSIFY OBJECTS
2y 10m to grant Granted Sep 08, 2026
Patent 12718340
METHOD FOR DETERMINING AREAS OF LAND COMPATIBLE WITH THE INSTALLATION OF PHOTOVOLTAIC PANELS
2y 10m to grant Granted Aug 25, 2026
Patent 12700132
INFORMATION PROCESSING SYSTEM, INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND RECORDING MEDIUM
3y 6m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
76%
Grant Probability
99%
With Interview (+27.2%)
3y 0m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 603 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month