Prosecution Insights
Last updated: October 04, 2026
Application No. 17/789,132

METHOD AND APPARATUS FOR TRAINING INFORMATION PREDICTION MODELS, METHOD AND APPARATUS FOR PREDICTING INFORMATION, AND STORAGE MEDIUM AND DEVICE THEREOF

Final Rejection §103
Filed
Jun 24, 2022
Priority
Dec 25, 2019 — CN 201911360658.2 +1 more
Examiner
MAUNI, HUMAIRA ZAHIN
Art Unit
2141
Tech Center
2100 — Computer Architecture & Software
Assignee
BIGO TECHNOLOGY PTE. LTD.
OA Round
4 (Final)
47%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
85%
With Interview

Examiner Intelligence

Grants 47% of resolved cases
47%
Career Allowance Rate
14 granted / 30 resolved
-8.3% vs TC avg
Strong +38% interview lift
Without
With
+38.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
25 currently pending
Career history
60
Total Applications
across all art units

Statute-Specific Performance

§101
33.4%
-6.6% vs TC avg
§103
50.7%
+10.7% vs TC avg
§102
1.7%
-38.3% vs TC avg
§112
14.2%
-25.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 30 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendments filed 06/19/2026 have been entered. Claims 1, 7, 9-11, and 14-17 remain pending within the application. The amendments filed 06/19/2026 are sufficient to overcome each and every objection previously set forth in the Non-Final Office Action mailed 03/19/2026. The objections have been withdrawn. The amendments filed 06/19/2026 are sufficient to overcome the 112(b) rejections previously set forth in the Non-Final Office Action mailed 03/19/2026. The rejections have been withdrawn. Claim Objections Claims 1, 10, and 15 are objected to because of the following informalities: “and an implicit vector output by the embedding layer are as an input of the fully connected layer” should be “and an implicit vector output by the embedding layer are used as an input to the fully connected layer”. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 7, 9-11, and 14-17 are rejected under 35 U.S.C. 103 as being unpatentable over Bilenko et al. (US 20140337096 A1), hereafter Bilenko, in view of Zhang et al. ("A Deep Joint Network for Session-based News Recommendations with Contextual Augmentation"), hereafter Zhang, in further view of Zeng et al. (US 20170161639 A1), hereafter Zeng. Regarding claim 1, Bilenko discloses: A method for recommending information, executed by a recommendation system, the method comprising: acquiring an information item to be recommended from an information prediction computer device, and recommending the information item to a user (Bilenko, Fig. 6, ¶[0057], ¶[0099] and ¶[0106] teaches acquiring and recommending content items, such as ads, to a user by a recommendation system), wherein a third information recommendation model is trained by a computer device configured for training an information recommendation model, and (Bilenko, Fig. 5 teaches a training system configured to train a final deployed model 518 as the third information prediction model), the computer device publishes the third information recommendation model to a server to enable the information prediction computer device to acquire the third information recommendation model from the server (Bilenko, Fig. 12, Fig. 13, Fig. 6 and ¶[0086-0087] teaches publishing the trained prediction model in the prediction module to a server), the information prediction computer device acquires the third information recommendation model and samples corresponding to candidate information items, determines prediction results corresponding to the candidate information items using the third information recommendation model, (Bilenko, Fig. 12, Fig. 5, Fig. 6 and ¶[0087], ¶[0094], and ¶[0106] teaches the recommended ad to be determined based on the prediction results corresponding to candidate information items, i.e. candidate ads, where the prediction results are generated by the system device using the final deployed model 518 as the third information prediction model), determines the information item to be recommended in the candidate information items based on the prediction results (Bilenko, ¶[0106] and Fig. 12 teaches determining the ad to be displayed based on the prediction result), wherein the third information recommendation model is periodically acquired by training by the computer device for training the information recommendation model through the following processes: acquiring a set of training samples corresponding to a current training period (Bilenko, Fig. 1, ¶[0032] and ¶[0057] teaches Data Collection Process 108 periodically acquiring a set of training samples corresponding to a current training period for training an information prediction model 106), wherein training samples in the set of training samples comprise feature items, feature attribute values corresponding to the feature items, and behavior data of a user for information items, the feature items comprising at least one of features of the user and features of the information items (Bilenko, Figs. 1 and 2, ¶[0002] and ¶[0033-0034] teaches feature vectors of feature items with corresponding attribute values, and user-related aspects as behavior data of a user for information items), a training period is measured according to time or according to the number of the training samples (Bilenko, Fig. 8 teaches the training period to be measured according to time), acquiring first behavior statistics amounts in first behavior statistics data in a first information recommendation model corresponding to feature attribute values present in the set of training samples (Bilenko, Fig. 7 and ¶[0090-0092] teaches updating the statistical information to acquire first behavior statistics amounts in the first behavior statistics data in a first information recommendation model corresponding to the feature attribute values present in the set of training samples), acquiring current behavior statistics amounts corresponding to the feature attribute values by superimposing behavior data corresponding to the feature attribute values present in the set of training samples … (Bilenko, Figs. 7 and 10, Fig. 11, ¶[0102-0104] teaches acquiring current behavior statistics amounts corresponding to the feature attribute values by superimposing behavior data on first behavior statistic amount through the prediction model providing plural instances of statistical information and using post-deployment data to perform further training), acquiring current behavior statistics data by aggregating the current behavior statistics amounts corresponding to the feature attribute values (Bilenko, Fig. 11 and ¶[0104] teaches generating subsets of data via aggregation module as aggregating the current behavior statistics amounts corresponding to the feature attribute values), wherein the first information recommendation model is an information recommendation model acquired by training in a previous training period, or in a case that the current training period is a first training period, a predetermined initialization information recommendation model is set as the first information recommendation model (Bilenko, ¶[0090-0092] teaches updating and acquiring recommendation models from previous training periods, where an initial training period produces a predetermined initialization information recommendation model), the first information recommendation model comprises an information recommendation model based on deep neural networks (DNN) (Bilenko, ¶[0082] and ¶[0098] teaches the prediction model to be a deep neural network), the first information recommendation model comprises …a fully connected layer, the fully connected layer receiving …the first behavior statistics data (Bilenko, ¶[0098] teaches the model to comprise fully connected layers that receive the first behavior statistics data through input), an input vector as an embedding layer received by the input layer, i.e. fully connected layer of a neural network of the prediction model), acquiring a second information recommendation model by replacing the first behavior statistics data in the first information recommendation model with the current behavior statistics data (Bilenko, Fig. 11 and ¶[0104] teaches generating a model, i.e. acquiring a second information recommendation model, by replacing behavior statistics data with current behavior statistics data through update), acquiring a trained third information recommendation model by updating parameters of … the fully connected layer in the second information recommendation model by means of training the second information recommendation model based on the set of training samples (Bilenko, Fig. 5, ¶[0082] and ¶[0098] teaches generating a trained third model 518 by updating parameters of the fully connected layer in the second information prediction model by means of training the second information prediction model during model training, based on the set of training samples in collected data 510), wherein the behavior statistics data in the second information recommendation model and an implicit vector … are as an input of the fully connected layer (Bilenko, ¶[0098] teaches behavior statistics data and implicit vectors to be input to the fully connected layers). Bilenko teaches acquiring first behavior statistics amounts, but does not teach: calculating a product of the first … amounts and a predetermined time decay factor. Zhang teaches: calculating a product of the first … amounts and a predetermined time decay factor (Zhang, page 206, left column, paragraph 2, last 2 lines “The decay rates are multiplied by the output values from LSTM RNN layer to form the final outputs” teaches calculating a product of the first amounts and a predetermined time decay factor). Bilenko teaches acquiring current behavior statistics amounts corresponding to the feature attribute values by superimposing behavior data corresponding to the feature attribute values present in the set of training samples … but does not teach superimposing the data on the product. Zhang teaches: superimposing … data … on the product (Zhang, page 206, Equation 15 and 2 lines below equation 15 “λ is the parameter that needs to be tuned during training, and controls the decay rate for the news.”, page 202, right column, paragraph 2, last 3 lines “we adopt time-decay function to reduce the weight of the historical news articles, and character-level encoding to alleviate sparsity problem.” And Fig. 1 teaches superimposing data on the product throughout training of the model recited in Figure 1). Bilenko and Zhang are analogous art because they are from the same field of endeavor, feature engineering, recommendations, and machine learning models. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Bilenko to include calculating a product of the first amounts and a predetermined time decay factor and superimposing data on the product, based on the teachings of Zhang. One of ordinary skill in the art would have been motivated to make this modification in order to improve accuracy while simplifying feature engineering steps, as suggested by Zhang (Zhang, page 203, left column, paragraph 1, last line). While Bilenko teaches the first information recommendation model comprises … a fully connected layer, the fully connected layer receiving …the first behavior statistics data, acquiring a trained third information recommendation model by updating parameters of … the fully connected layer in the second information recommendation model by means of training the second information recommendation model based on the set of training samples, and wherein the behavior statistics data in the second information recommendation model and an implicit vector … are as an input of the fully connected layer, they do not explicitly teach an embedding layer preceding the fully connected layer. Zhang discloses: … model comprises an embedding layer and a fully connected layer, the fully connected layer receiving the embedding layer… (Zhang, Figure 1 and the paragraph below Figure 1, lines 3-5 “After character-level embedding into matrix as input, a 2-layer convolutional neural network is deployed” teaches a model comprises an embedding layer and a fully connected layer, the fully connected layer receiving the embedding layer), …updating parameters of the embedding layer… (Zhang, Figure 1 and Table 2 teaches training and updating the parameters of the embedding layer), …an implicit vector output by the embedding layer are as an input of the fully connected layer… (Zhang, Figure 1 and equations 1-3 teaches the implicit vectors output by the embedding layers to be input to the fully connected layer). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Bilenko to include model comprises an embedding layer and a fully connected layer, the fully connected layer receiving the embedding layer, updating parameters of the embedding layer, and an implicit vector output by the embedding layer are as an input of the fully connected layer, based on the teachings of Zhang. One of ordinary skill in the art would have been motivated to make this modification in order to improve accuracy while simplifying feature engineering steps, as suggested by Zhang (Zhang, page 203, left column, paragraph 1, last line). Bilenko, in view of Zhang, does not disclose: instructing a storage device storing the set of training samples to delete the set of training samples corresponding to the current training period upon completion of the training. Zeng discloses: instructing a storage device storing the set of training samples to delete the set of training samples corresponding to the current training period upon completion of the training (Zeng, ¶[0069] teaches deleting training samples upon completion of training for a training period). Bilenko and Zhang are analogous art because they are from the same field of endeavor, feature engineering, recommendations, and machine learning models. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Bilenko, in view of Zhang, to include instructing a storage device storing the set of training samples to delete the set of training samples corresponding to the current training period upon completion of the training, based on the teachings of Zeng. One of ordinary skill in the art would have been motivated to make this modification for more accurate recommendations while maintaining efficiency, as suggested by Zeng (¶[0034]). Regarding claim 7, Bilenko, in view of Zhang, in further view of Zeng, discloses the method according to claim 1. Bilenko further discloses: wherein the feature attribute values are represented by hash values (Bilenko, ¶[0051] teaches feature attribute values represented by hash values). Regarding claim 9, Bilenko, in view of Zhang, in further view of Zeng, discloses the method according to claim 1. Bilenko further discloses: wherein the first (Bilenko, ¶[0049] teaches an information recommendation model based on click through rates). Claim 10 is substantially similar to claim 1, and thus are rejected on the same basis as claim 1. Regarding claim 11, Bilenko, in view of Zhang, in further view of Zeng, discloses the method according to claim 10. Bilenko further discloses: wherein the information recommendation model comprises an information recommendation model based on click through rates CTR (Bilenko, ¶[0049] teaches an information recommendation model based on click through rates), determining, based on the output result of the information recommendation model, the prediction result corresponding to the candidate information items comprises: determining, based on the output result of the information recommendation model, a CTR prediction result corresponding to the candidate information items (Bilenko, ¶[0048-0049], ¶[0067-0069] teaches a training system to determine a CTR prediction result corresponding to candidate information, based on the output result of the prediction model, by forming clusters of user IDs that have similar click through rates in a manner that minimizes the loss of the predictive accuracy), upon determining, based on the output result of the information recommendation model, the prediction results corresponding to the candidate information items, the method further comprises: determining … the candidate information items based on the CTR prediction result (Bilenko, ¶[0001] teaches serving one or more ads having high click probabilities based on the model output as determining the candidate information items based on the CTR prediction result), determining… an information item to be recommended in the candidate information items (Bilenko, ¶[0106] teaches determining which ads to display to users as determining an information item to be recommended in the candidate information items). While Bilenko teaches determining … the candidate information items based on the CTR prediction result, and determining… an information item to be recommended in the candidate information items, they do not explicitly disclose determining the order of information items. Zhang teaches: determining an order of … information items (Zhang, page 203, paragraph 1, last 4 lines “…the system is to predict…A recommendation … is an ordered list of recommended items, where we would want to see the next item as close to the top as possible.” Teaches determining an ordered list of recommendations as determining an order of information items). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Bilenko to include determining an order of … information items, based on the teachings of Zhang. One of ordinary skill in the art would have been motivated to make this modification in order to improve accuracy while simplify feature engineering steps, as suggested by Zhang (Zhang, page 203, left column, paragraph 1, last line). Regarding claim 14, Bilenko, in view of Zhang, in further view of Zeng, discloses the method for recommending information as defined in claim 1. Bilenko further discloses: A non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when run by a processor, causes the processor to perform the method for recommending information as defined in claim 1 (Bilenko, Fig. 13 and ¶[0109-0114] teaches a computer-readable storage medium, storing a computer program, wherein the computer program, when run by a processor, causes the processor to perform the method for training information recommendation models as defined in claim 1). Claim 15 is substantially similar to claim 1, and thus is rejected on the same basis as claim 1. Regarding claim 16, Bilenko, in view of Zhang, in further view of Zeng, discloses the method for recommending information as defined in claim 10. Bilenko further discloses: A computer device for predicting information, comprising: a memory, a processor, and a computer program that is stored in the memory and runnable in the processor, wherein the processor, when running the computer program, is caused to perform the method for recommending information as defined in claim 10 (Bilenko, Fig. 13 and ¶[0109-0114] teaches a computer device for predicting information, comprising: a memory, a processor, and a computer program that is stored in the memory and runnable in the processor, wherein the processor, when running the computer program, is caused to perform the method for recommending information as defined in claim 10). Regarding claim 17, Bilenko, in view of Zhang, in further view of Zeng, discloses the method for recommending information as defined in claim 10. Bilenko further discloses: A non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when run by a processor, causes the processor to perform the method for recommending information as defined in claim 10 (Bilenko, Fig. 13 and ¶[0109-0114] teaches a computer-readable storage medium, storing a computer program, wherein the computer program, when run by a processor, causes the processor to perform the method for predicting information as defined in claim 10). Response to Arguments Applicant's arguments filed 06/19/2026 have been fully considered with regards to the 35 U.S.C. 102/103 rejection, but they are not persuasive. The applicant asserts on page 17 of the remarks “It can be seen that the updating of the model in Bilenko is based on the master dataset, which includes the data acquired before and after the model deployment, in contrast to claim 1 wherein the model update is achieved only based on the data of the current training period.”. In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., updating a model without post deployment data) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). The independent claims recite “acquiring a set of training samples corresponding to a current training period”, the BRI of which includes any data received during a current training cycle, and thus is taught by Bilenko (Fig. 1 and Fig. 7). The applicant asserts on page 19 of the remarks “Bilenko also fails to disclose the layer structure of the recommendation model and the function or role of the behavior statistics data in the layer structure of the recommendation model. Zhang does not disclose or render obvious elements of claim 1 whether taken alone or in combination with Bilenko.” The examiner respectfully disagrees, as Fig. 11 and ¶[0098] teaches the model generation and the function or role of the behavior statistics data in the layer structure of the recommendation model, where feature information for input layers of the generated model rely on behavior statistics data. The new ground of rejection also relies on the Zhang reference to teach elements of the layer structure of the model (Zhang, Figure 1). The examiner refers to the rejection under 35 USC § 103 for claim 1 in the current office action for more details. Claims 10 and 15 and substantially similar to claim 1, and thus are rejected on the same basis. Claims dependent on the independent claims do not overcome the deficiencies of the rejected independent claims. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUMAIRA ZAHIN MAUNI whose telephone number is (703)756-5654. The examiner can normally be reached Monday - Friday, 9 am - 5 pm (ET). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MATT ELL can be reached at (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /H.Z.M./Examiner, Art Unit 2141 /MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Show 2 earlier events
Sep 04, 2025
Response Filed
Oct 09, 2025
Final Rejection mailed — §103
Dec 09, 2025
Response after Non-Final Action
Jan 06, 2026
Request for Continued Examination
Jan 14, 2026
Response after Non-Final Action
Mar 19, 2026
Non-Final Rejection mailed — §103
Jun 19, 2026
Response Filed
Sep 16, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711418
VARIABLE-OUTPUT-SPACE PREDICTION MACHINE LEARNING MODELS USING CONTEXTUAL INPUT EMBEDDINGS
4y 3m to grant Granted Aug 18, 2026
Patent 12705309
COMPUTER-IMPLEMENTED DETECTION METHOD, NON-TRANSITORY COMPUTER-READABLE RECORDING MEDIUM, AND COMPUTING SYSTEM
4y 5m to grant Granted Aug 11, 2026
Patent 12688409
OPTIMIZING SEND TIME FOR ELECTRONIC COMMUNICATIONS
5y 5m to grant Granted Jul 21, 2026
Patent 12682253
METHOD AND DEVICE FOR CONSTRUCTING DECISION TREE
4y 8m to grant Granted Jul 14, 2026
Patent 12670385
Technique for Retraining Operational Neural Networks Using Synthetically Generated Retraining Data
4y 3m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
47%
Grant Probability
85%
With Interview (+38.1%)
4y 1m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 30 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month