Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This action is responsive to the following communication: Amendment filed Jul. 16, 2026. This Action is made Final.
Claims 1-20 are pending in the case. Claims 1, 9 and 16 are independent claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-17 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Rivlin et al. (hereinafter Rivlin) U.S. Patent Publication No. 2019/0043490 in view of Nobari et al. (hereinafter Nobari) “Query Understanding via Entity Attribute Identification” (hereinafter Rivlin) 2018.
With respect to independent claim 1, Rivlin teaches a system for machine learning based updates of domain specific language (DSL) parameters, the system comprising: one or more memories; and one or more processors, communicatively coupled to the one or more memories (see e.g., Fig. 2 Para [31]-“the method 400 may be implemented in one or more modules as a set of logic instructions stored in a machine- or computer-readable storage medium such as random access memory (RAM), read only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc., in configurable logic such as, for example, programmable logic arrays (PLAs), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and fixed-functionality logic hardware using circuit technology such as, for example, application specific integrated circuit (ASIC), complementary metal oxide semiconductor (CMOS) or transistor-transistor logic (TTL) technology, or any combination thereof.”), configured to:
obtain audio data from one or more individuals, the audio data indicating information associated with one or more attributes associated with an entity (see e.g., Para [33]-[35] - “the speech may be captured as audio via in-person conversations, phone conversations, oral readings of passages, and other means of capturing speech from the user. In another embodiment, the written word of the user may be captured via emails, texts, notes, and other documents written by the user. In yet another embodiment, both audio and the written word of the user may be captured.”);
analyze the audio data to identify one or more words or phrases from the audio data (see e.g., Para [38]-[49]- “a prioritized n-best list of the n most probable words is created. The n-best list may be created using any of the prediction engines listed above. In one embodiment, an n-gram model may be used to predict the n-best list. For example, returning to the stutter example in which a user might state the following, “I remember when I was a kid playing in the neighborhood, we used to play ba-ba-ba-<pause> . . . . An example prioritized n-best list, based on the stutter sound of ba-ba-ba, the words spoken prior to the stutter (“I remember when I was a kid playing in the neighborhood, we used to play”),”);
detect one or more keywords, from the one or more words or phrases, indicative of the entity (see e.g., Para [30]-[38]-“they are more likely to experience memory lapses where they cannot readily remember names, places and objects (i.e., things) … For example, a user might say the following, “I can't find my <long pause> . . . . A stutter may be detected when the user repeats a portion of a word several times (the repetition of a sound several times) or uses a pause word like, for example, “uhhhhhh”. For example, a user might state the following, “I remember when I was a kid playing in the neighborhood, we used to play ba-ba-ba-<pause> . . . . The process then proceeds to block 506 … In block 506, the missing or stuttered word is predicted. Predicting the missing or stuttered word is further described below with reference to FIG. 6. The process then proceeds to block 508.”); provide, to a machine learning model, an indication of the entity and the one or more words or phrases (see e.g., Para [16][27][30]- “The personal language model evolves and adapts over time, and can be used in real time to retain all aspects of auditory and cognitive processing capabilities of a user … In one embodiment, the speaking aid system 200 may also be used to build the personal speech model as well as adapt or adjust the model based on interactions with the system 200.”);
obtain, from the machine learning model and based on providing the indication of the entity and the one or more words or phrases, an indication of one or more DSL parameters that are to be updated for the entity (see e.g., Para [30][33][45]-[48]- “the words used prior to a missing word or a stuttered word are input into the prediction engine. The process then proceeds to block 560” The examiner notes that the personal language models correspond to the recited “DSL”” The prediction engine may be a neural network, a Hidden Markov Model (HMM), a statistical language model, or any other prediction engine that may predi”),
wherein the one or more DSL parameters are associated with the one or more attributes; and update, based on an output of the machine learning model, the one or more DSL parameters to obtain one or more updated DSL parameters (see e.g., Para [50][51] - “In one embodiment, only a single word may be predicted at a time. For example, returning to the missing word example in which a user might say “I can't find my <long pause> . . . ”, the next word is predicted using the words spoken prior to the stutter (“I can't find my”), the personal language model for the user, and the current context data for the user. In the personal language model for the user, there are several utterances by the user in which the user has misplaced their mobile phone. The word “phone” would be an excellent prediction to whisper or display to the user in this instance …Embodiments are not limited by the number of predictions that may be used. One or more predictions may be made without departing from the scope.”).
Rivlin does not expressly show the audio data indicating information associated with one or more attributes associated with an entity, wherein the entity is a commercial entity. However, Nobari discloses similar feature (see e.g. Page 1759 – 1761 - “Extracting entity attributes of queries is beneficial for answering the queries in tasks such as question answering and entity retrieval. It has been shown that joint entity linking and attribute identification of queries improves question answering over knowledge bases [16, 17]. Similarly, entity retrieval approaches can benefit from entity attribute identification by having a focused selection of entity attributes [7] and using them to build fielded representation of entities [9]. Entity attribute identification can be also employed in the e-commerce websites to improve search results and boost sites’ advertising profits and recommendation quality.”) Both Rivlin and Nobari are directed to user input data analysis methods. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Rivlin and Nobari in front of them to modify the system of Rivlin to include the above feature. The motivation to combine Rivlin and Nobari comes from Nobari. Nobari discloses the motivation to improve query response by identify entity attributes (see e.g. Page 1759 – “Extracting entity attributes of queries is beneficial for answering the queries in tasks such as question answering and entity retrieval. It has been shown that joint entity linking and attribute identifi cation of queries improves question answering over knowledge bases [16, 17]. Similarly, entity retrieval approaches can benefit from entity attribute identification by having a focused selection of entity attributes [7] and using them to build fielded representation of entities [9]. Entity attribute identification can be also employed in the e-commerce websites to improve search results and boost sites’ advertising profits and recommendation quality. Consider, for example, the query “nike shoes size 38”, where the attribute size can be used to filter out irrelevant products or advertising similar products from other brands. Motivated by the above reasons, we set out to focus on identify ing entity attributes that help answering a query.”). This motivation for combination also applies to the remaining claims which depend on this combination.
With respect to dependent claim 2, the modified Rivlin teaches the one or more DSL parameters are associated with an application or another machine learning model that is capable of determining attributes of entities from raw data (see e.g., Para [27][30][35]-[43] – “The personal language model evolves and adapts over time, and can be used in real time to retain all aspects of auditory and cognitive processing capabilities of a user … In one embodiment, the speaking aid system 200 may also be used to build the personal speech model as well as adapt or adjust the model based on interactions with the system 200.” – The personal language model is a machine learning model. The raw data corresponds to the captured audio data.” the speaking aid application corresponds to the recited “application.”).
With respect to dependent claim 3, the modified Rivlin teaches provide, to an application or another machine learning model, an indication of the one or more updated DSL parameters (see e.g., Para [29]-[35]- the speaking aid application corresponds to the recited “application.”); provide, to the application or the other machine learning model, transaction data (see e.g., Para [29]); and obtain, from the application or the other machine learning model, an indication of attributes of one or more entities associated with the transaction data based on providing the indication of the one or more updated DSL parameters (see e.g., 5A 5B Para [36]-[51]).
With respect to dependent claim 4, the modified Rivlin teaches the one or more processors, to provide the indication of the entity and the one or more words or phrases, are configured to: provide, to the machine learning model, an indication of one or more contextual inputs associated with the audio data (see e.g., Para [29]-[34] -“context metadata may be added to make finding and working with particular instances of recordings easier. For example, a recording may include information and context metadata that includes, but is not limited to, user name, date and time created, who the user is talking to or sending a message to, what the user is talking or writing about, location of the user, activity of the user”).
With respect to independent claim 5, the modified Rivlin teaches the one or more contextual inputs include at least one of: location information, a user identifier associated with the audio data, an account record of a user associated with the audio data, or an image associated with the entity (see e.g., Para [29]-[34] -“context metadata may be added to make finding and working with particular instances of recordings easier. For example, a recording may include information and context metadata that includes, but is not limited to, user name, date and time created, who the user is talking to or sending a message to, what the user is talking or writing about, location of the user, activity of the user”).
With respect to independent claim 6, the modified Rivlin teaches the one or more processors, to update the one or more DSL parameters, are configured to: provide, to a device, an indication of the output of the machine learning model; and obtain, via a user input and based on providing the indication of the output, an indication of the one or more updated DSL parameters (see e.g., 5A 5B Para [36]-[51] - “a prioritized n-best list of the n most probable words is created. The n-best list may be created using any of the prediction engines listed above. In one embodiment, an n-gram model may be used to predict the n-best list. … only a single word may be predicted at a time. For example, returning to the missing word example in which a user might say “I can't find my <long pause> . . . ”, the next word is predicted using the words spoken prior to the stutter (“I can't find my”), the personal language model for the user, and the current context data for the user. ”).
With respect to dependent claim 7, the modified Rivlin teaches the one or more DSL parameters include an entity relationship parameter indicative of a relationship between the entity and one or more other entities (see e.g., Para [33]-[34][75] [77]-“ the personal language model for the user is generated in a continuous manner over time by monitoring conversations of the user at different contexts and storing the conversations with metadata describing the different contexts to create a timeline of information about the user.”).
With respect to dependent claim 8, the modified Rivlin teaches the one or more DSL parameters include metadata associated with the entity (see e.g., Para [38][75] – “the personal language model for the user is generated in a continuous manner over time by monitoring conversations of the user at different contexts and storing the conversations with metadata describing the different contexts to create a timeline of information about the user.”).
With respect to independent claim 9, Rivlin teaches a method of machine learning based updates of parameters, comprising: obtaining, by a device, an input indicating information associated with one or more attributes of an entity; analyzing, by the device, the input to identify one or more words or phrases from the input (see e.g., Para [33]-[35] - “the speech may be captured as audio via in-person conversations, phone conversations, oral readings of passages, and other means of capturing speech from the user. In another embodiment, the written word of the user may be captured via emails, texts, notes, and other documents written by the user. In yet another embodiment, both audio and the written word of the user may be captured.”); detecting, by the device, one or more domain specific language (DSL) keywords, from the one or more words or phrases, indicative of one or more parameters associated with the entity (see e.g., Para [30]-[38]-“they are more likely to experience memory lapses where they cannot readily remember names, places and objects (i.e., things) … For example, a user might say the following, “I can't find my <long pause> . . . . A stutter may be detected when the user repeats a portion of a word several times (the repetition of a sound several times) or uses a pause word like, for example, “uhhhhhh”. For example, a user might state the following, “I remember when I was a kid playing in the neighborhood, we used to play ba-ba-ba-<pause> . . . . The process then proceeds to block 506 … In block 506, the missing or stuttered word is predicted. Predicting the missing or stuttered word is further described below with reference to FIG. 6. The process then proceeds to block 508.”); providing, by the device and to a machine learning model, an indication of the entity and the one or more words or phrases (see e.g., Para [16][27][30]- “The personal language model evolves and adapts over time, and can be used in real time to retain all aspects of auditory and cognitive processing capabilities of a user … In one embodiment, the speaking aid system 200 may also be used to build the personal speech model as well as adapt or adjust the model based on interactions with the system 200.”); obtaining, by the device and via an output of the machine learning model and based on providing the indication of the entity and the one or more words or phrases, an indication of updates to be made for the one or more parameters (see e.g., Para [30][33][45]-[48]- “the words used prior to a missing word or a stuttered word are input into the prediction engine. The process then proceeds to block 560” The examiner notes that the personal language models correspond to the recited “DSL”” The prediction engine may be a neural network, a Hidden Markov Model (HMM), a statistical language model, or any other prediction engine that may predict”); and performing, by the device, an action based on the output of the machine learning model (see e.g., Para [50][51] - “In one embodiment, only a single word may be predicted at a time. For example, returning to the missing word example in which a user might say “I can't find my <long pause> . . . ”, the next word is predicted using the words spoken prior to the stutter (“I can't find my”), the personal language model for the user, and the current context data for the user. In the personal language model for the user, there are several utterances by the user in which the user has misplaced their mobile phone. The word “phone” would be an excellent prediction to whisper or display to the user in this instance …Embodiments are not limited by the number of predictions that may be used. One or more predictions may be made without departing from the scope.”).
Rivlin does not expressly show the parameters are brand parameters. However, Rivlin is directed to words recognition in general and does not limit itself to any specific subject matter that words represent. Therefore, it would have been obvious to include brand parameters.
Rivlin does not expressly show the audio data indicating information associated with one or more attributes associated with an entity, wherein the entity is a commercial entity. However, Nobari discloses similar feature (see e.g. Page 1759 – 1761 - “Extracting entity attributes of queries is beneficial for answering the queries in tasks such as question answering and entity retrieval. It has been shown that joint entity linking and attribute identification of queries improves question answering over knowledge bases [16, 17]. Similarly, entity retrieval approaches can benefit from entity attribute identification by having a focused selection of entity attributes [7] and using them to build fielded representation of entities [9]. Entity attribute identification can be also employed in the e-commerce websites to improve search results and boost sites’ advertising profits and recommendation quality.”) Both Rivlin and Nobari are directed to user input data analysis methods. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Rivlin and Nobari in front of them to modify the system of Rivlin to include the above feature. The motivation to combine Rivlin and Nobari comes from Nobari. Nobari discloses the motivation to improve query response by identify entity attributes (see e.g. Page 1759 – “Extracting entity attributes of queries is beneficial for answering the queries in tasks such as question answering and entity retrieval. It has been shown that joint entity linking and attribute identifi cation of queries improves question answering over knowledge bases [16, 17]. Similarly, entity retrieval approaches can benefit from entity attribute identification by having a focused selection of entity attributes [7] and using them to build fielded representation of entities [9]. Entity attribute identification can be also employed in the e-commerce websites to improve search results and boost sites’ advertising profits and recommendation quality. Consider, for example, the query “nike shoes size 38”, where the attribute size can be used to filter out irrelevant products or advertising similar products from other brands. Motivated by the above reasons, we set out to focus on identify ing entity attributes that help answering a query.”). This motivation for combination also applies to the remaining claims which depend on this combination.
With respect to dependent claim 10, Rivlin the modified teaches the input includes at least one of audio data or text data (see e.g., Para [33][34] -“both audio and the written word of the user may be captured”).
With respect to dependent claim 11, the modified Rivlin teaches performing the action comprises: updating, in a database, the one or more brand parameters based on the output of the machine learning model (see e.g., Para [35][36] – “store the personal language model in the cloud”).
With respect to dependent claim 12, the modified Rivlin teaches performing the action comprises: providing, to another device, an indication of the updates to be made for the one or more brand parameters (see e.g., Para [19][29]-[31] – Rivlin does not expressly show this feature. However, based on the teachings of Rivlin, it would have been obvious to include this feature because the updates are stored on a cloud server and any device that rely on the cloud service will have updated parameters).
With respect to dependent claim 13, the modified Rivlin teaches the input includes voice data or text data that is indicative of the updates (see e.g., Para [29]-[35] - “the speech may be captured as audio via in-person conversations, phone conversations, oral readings of passages, and other means of capturing speech from the user. In another embodiment, the written word of the user may be captured via emails, texts, notes, and other documents written by the user. In yet another embodiment, both audio and the written word of the user may be captured.”).
With respect to dependent claim 14, the modified Rivlin teaches the one or more brand parameters include an entity relationship parameter, and wherein performing the action comprises: updating an entity relationship graph to indicate a relationship between the entity and one or more other entities based on the entity relationship parameter (see e.g., Para [48] Claim 5– the examiner notes that it is well-known in the art that a neural network provides an entity relationship graph).
With respect to dependent claim 15, the modified Rivlin teaches the one or more brand parameters include at least one of: a brand name, one or more uniform resource locator addresses, one or more location addresses, one or more phone numbers, one or more email addresses, a merchant category code, one or more national parameters, one or more location-specific parameters, or a relationship parameter indicating a relationship between two or more brand parameters (see e.g., Para [29] – “a recording may include information and context metadata that includes, but is not limited to, user name, date and time created, who the user is talking to or sending a message to, what the user is talking or writing about, location of the user, activity of the user, and any other information that may describe the content of the recording or document. The information collected may be used to service the user 302 for the benefit of the user 302 in the near or distant future when the user 302 needs help trying to speak a missing or stuttered word. So, over years of time the system is learning the person's language model by listening to a plurality of in-person conversations, phone conversations, or other means of communication, such as, for example, by storing texts, emails, notes, etc.”).
Claim 16 is rejected for the similar reasons discussed above with respect to claim 1.
With respect to dependent claim 17, the modified Rivlin teaches the one or more DSL parameters are associated with determining attributes of entities from transaction data (see e.g., Para [29] – “a recording may include information and context metadata that includes, but is not limited to, user name, date and time created, who the user is talking to or sending a message to, what the user is talking or writing about, location of the user, activity of the user, and any other information that may describe the content of the recording or document. The information collected may be used to service the user 302 for the benefit of the user 302 in the near or distant future when the user 302 needs help trying to speak a missing or stuttered word. So, over years of time the system is learning the person's language model by listening to a plurality of in-person conversations, phone conversations, or other means of communication, such as, for example, by storing texts, emails, notes, etc.”).
With respect to dependent claim 19, the modified Rivlin teaches the one or more instructions, that cause the device to provide the indication of the entity and the one or more words or phrases, cause the device to: provide, to the machine learning model, an indication of one or more contextual inputs associated with the audio data (see e.g., Para [29]-[34] -“context metadata may be added to make finding and working with particular instances of recordings easier. For example, a recording may include information and context metadata that includes, but is not limited to, user name, date and time created, who the user is talking to or sending a message to, what the user is talking or writing about, location of the user, activity of the user”).
With respect to dependent claim 20, the modified Rivlin teaches the one or more instructions, that cause the device to update the one or more DSL parameters, cause the device to: provide, to another device, an indication of the output of the machine learning model (see e.g., Para [19][29]-[31] – Rivlin does not expressly show this feature. However, based on the teachings of Rivlin, it would have been obvious to include this feature because the updates are stored on a cloud server and any device that rely on the cloud service will have updated parameters); and obtain, via a user input and based on providing the indication of the output, an indication of the one or more updated DSL parameters (see e.g., 5A 5B Para [36]-[51] - “a prioritized n-best list of the n most probable words is created. The n-best list may be created using any of the prediction engines listed above. In one embodiment, an n-gram model may be used to predict the n-best list. … only a single word may be predicted at a time. For example, returning to the missing word example in which a user might say “I can't find my <long pause> . . . ”, the next word is predicted using the words spoken prior to the stutter (“I can't find my”), the personal language model for the user, and the current context data for the user. ”).
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Rivlin in view of Nobari and further in view of Oros (hereinafter Oros) U.S. Patent Publication No. 2020/0134374.
With respect to dependent claim 18, Rivlin does not expressly show provide, to another machine learning model that uses the one or more updated DSL parameters, transaction data; and obtain, from the other machine learning model, an indication of attributes of one or more entities associated with the transaction data based on providing the indication of the one or more updated DSL parameters.
However, Oros teaches the similar feature (see e.g. para [29][36][52]). Both Rivlin and Oros are directed to machine learning methods. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Rivlin and Oros in front of them to further modify the modified system of Rivlin to include the above feature. The motivation to combine Rivlin and Oros comes from Oros. Oros discloses the motivation to dynamically update a machine learning model based on any possible predefined trigger so that the machine learning model can perform according to most current update(see e.g. para [2]-[6]).
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). Further, a reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill the art, including nonpreferred embodiments. Merck & Co. v. Biocraft Laboratories, 874 F.2d 804, 10 USPQ2d 1843 (Fed. Cir.), cert. denied, 493 U.S. 975 (1989). See also Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); Celeritas Technologies Ltd. v. Rockwell International Corp., 150 F.3d 1354, 1361, 47 USPQ2d 1516, 1522-23 (Fed. Cir. 1998).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEIYONG WENG whose telephone number is (571)270-1660. The examiner can normally be reached on Mon.-Fri. 8 am to 5 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Matthew Ell, can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://portal.uspto.gov/external/portal. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/PEI YONG WENG/Primary Examiner, Art Unit 2141